Published July 18, 2024 | Version 2023.1.8

The OpenITI Millionaires

  • 1. ROR icon Aga Khan University

Description

This data set pertains to the largest works in the OpenITI corpus at or prior to 1000 AH and is based on the 2023.1.8 release of the corpus and the corresponding text reuse data between the books in the corpus, which is generated by running using passim on the corpus. 

We wanted to understand the extent to which a small number of persons produced a substantial percentage of the OpenITI corpus, on a word-count basis. We call the authors with work(s) over a million words the ‘millionaires’.

The data will be analysed in forthcoming publications by the KITAB project team, including a monograph by Sarah Bowen Savant under contract with Edinburgh University Press. KITAB is funded by the European Research Council under the European Union’s Horizon 2020 research and innovation programme, awarded to the KITAB project (Grant Agreement No. 772989, PI Sarah Bowen Savant), hosted at Aga Khan University, London. In addition, it has received funding from the Qatar National Library to aid in the adaptation of the passim algorithm for Arabic.

KITAB’s text reuse data is published on Zenodo and each version is the output of a separate run. The version number of each release corresponds to the corpus releases.

 

Files

OpenITI-Millionaires_ReleaseNotes_v2023.1.8.pdf

Files (172.0 MB)

Name Size Download all
md5:f03211590327e60faf788a50ae2bdf58
3.7 kB Download
md5:6c0cf8c92e2c319c8e21e87e1ae046e4
46.7 kB Preview Download
md5:58ea5e28f7b6a3b2252dd64186d9fb82
7.4 MB Download
md5:43c072bbf7ef1a273359b5a5c566e2e2
164.5 MB Download

Additional details

Funding

European Commission
KITAB - Exploring Cultural Memory in the Pre-Modern Islamic World (700–1500): Knowledge, Information Technology, and the Arabic Book 772989