Published May 20, 2025
| Version 1.0.0
Dataset
Open
Agapet (Christian Arabic HTR Model) Training Datasets
Authors/Creators
Description
These files consist of images and gold-standard, expert-corrected segmentation and transcription of Christian Arabic manuscripts in PAGE .XML format. These images were used for the training and testing of the Agapet HTR models for Christian Arabic hands, available in formats compatible with Transkribus and eScriptorium/Kraken. The specific contents of the dataset are as follows:
- Full segmentation and transcription of SA-418, a 13th-century manuscript, totalling 342 pages.
- Full segmentation and transcription of Sinai Arabic 423, a 17th-century manuscript, totalling 478 pages across 239 images
- 11 segmented and transcribed pages of BnF Arabe 76, a 14th-century manuscript, used for further testing of the eScriptorium models
These datasets are released to allow independent training, testing, and development of models for computational recognition and analysis of Christian Arabic hands.
Files
BnF Arabe 76 pp 68-78.zip
Additional details
Related works
- Is supplement to
- 10.5281/zenodo.14382311 (DOI)
Dates
- Submitted
-
2025-05-20
Software
- Programming language
- XML