Published August 11, 2026
| Version v1
Dataset
Open
NOMOCRAT Maltese OCR data set
Authors/Creators
Description
Data set produced by the NOMOCRAT project, a project with the aim to create a visual text extraction model from Maltese language PDFs using layout analysis, optical character recognition, and document reading order determination. The pages were taken from public Maltese language PDFs that can be found in the dokumenti.mt repository. Size is too small to be used for training but can be used for evaluation.