OpenITI MAKHZAN

Parkes Allen, Jonathan; Mullan, John; Nigst, Lorenz; Barber, Mathew; Shahid Khan, Taimoor; Seydi, Masoumeh; Chen, Danlu; Weng, Yufei; Vogler, Nikolai; Murel, Jacob; Eshera, Osama; Berg-Kirkpatrick, Taylor; Smith, David; Bowen Savant, Sarah; Thomas Miller, Matthew; Verkinderen, Peter

doi:10.5281/zenodo.16410106

Published November 7, 2025 | Version 2025.1.1

Dataset Open

OpenITI MAKHZAN

1. University of Maryland, College Park
2. The Aga Khan University (International) in the United Kingdom
3. University of California, San Diego
4. Northeastern University

OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript Data

The Open Islamicate Texts Initiative (OpenITI) was founded in 2017 with the goal of building the digital infrastructure for the study of the premodern Islamic world. The underperformance of automatic transcription via Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) for Arabic-script languages long stymied the production of large-scale corpora in Islamicate languages, such as Persian, Arabic, Ottoman Turkish, and Urdu.

The OpenITI MAKHZAN dataset is a large aggregation of Arabic-script ground truth and evaluation data drawn from a wide variety of Persian, Arabic, Ottoman Turkish, and Urdu print and handwritten (manuscript) documents.

Note on Release Numbering: Version 2025.1.2—where 2025 is the year of the release, the first dotted number—.1—is the ordinal release number in 2025, and the second dotted number—.2—is the overall release number; the first dotted number will reset every year, while the second one will continue on increasing.

Files

OpenITI-Makhzan_Data_2025-1-1.zip

Files (4.1 GB)

Name	Size	Download all
OpenITI-Makhzan_Data_2025-1-1.zip md5:d1a75d748e74a28c09232f85aee5a0e5	4.1 GB	Preview Download
OpenITI-Makhzan_Metadata_2025-1-1.tsv md5:cd565d599ac58cbcf40d529c767faa4c	343.0 kB	Download
OpenITI-Makhzan_ReleaseNotes_2025-1-1.pdf md5:fd4ad7d140fb6fef82d60b7b6ae0fefe	117.9 kB	Preview Download

	All versions	This version
Views	282	164
Downloads	184	104
Data volume	398.2 GB	190.5 GB

OpenITI MAKHZAN

Authors/Creators

Description

Files

OpenITI-Makhzan_Data_2025-1-1.zip

Files (4.1 GB)