Published May 20, 2025 | Version 1.0.0

Agapet (Christian Arabic HTR Model) Training Datasets

  • 1. ROR icon University of Tübingen
  • 2. EDMO icon Harvard University

Description

These files consist of images and gold-standard, expert-corrected segmentation and transcription of Christian Arabic manuscripts in PAGE .XML format. These images were used for the training and testing of the Agapet HTR models for Christian Arabic hands, available in formats compatible with Transkribus and eScriptorium/Kraken. The specific contents of the dataset are as follows:

  • Full segmentation and transcription of SA-418, a 13th-century manuscript, totalling 342 pages.
  • Full segmentation and transcription of Sinai Arabic 423, a 17th-century manuscript, totalling 478 pages across 239 images
  • 11 segmented and transcribed pages of BnF Arabe 76, a 14th-century manuscript, used for further testing of the eScriptorium models

These datasets are released to allow independent training, testing, and development of models for computational recognition and analysis of Christian Arabic hands.

Files

BnF Arabe 76 pp 68-78.zip

Files (526.6 MB)

Name Size
md5:23348293f23da17ba20956971cb076b4
97.0 MB Preview Download
md5:4cf6ebb1be75511f59bd6c23726cfef0
366.1 MB Preview Download
md5:c27b5ed444d587f27f330dc7684dce85
63.5 MB Preview Download

Additional details

Related works

Is supplement to
10.5281/zenodo.14382311 (DOI)

Dates

Submitted
2025-05-20

Software

Programming language
XML