Published August 11, 2026 | Version v1

NOMOCRAT Maltese OCR data set

  • 1. ROR icon University of Malta
  • 2. University of Malta Faculty of Engineering
  • 3. University of Malta, Faculty of Engineering
  • 4. European Commission

Description

Data set produced by the NOMOCRAT project, a project with the aim to create a visual text extraction model from Maltese language PDFs using layout analysis, optical character recognition, and document reading order determination. The pages were taken from public Maltese language PDFs that can be found in the dokumenti.mt repository. Size is too small to be used for training but can be used for evaluation.

Files

layout_data.json

Files (29.6 MB)

Name Size Download all
md5:6b60a77ac94afbf71e7e26170a8d9126
667.4 kB Preview Download
md5:083ca8c492b52853e09e03df721964d2
541.2 kB Preview Download
md5:38e61c18c135f728602718ef96f64b00
28.1 MB Preview Download
md5:21a0447a17109f1b89b88e012a3bf01c
329.0 kB Preview Download

Additional details

Funding

Xjenza Malta
Research Excellence Programme REP-2024-057