There is a newer version of the record available.

Published December 19, 2020 | Version 2020.1.2

KITAB Text Reuse Data

  • 1. Northeastern University
  • 2. Aga Khan University
  • 3. University of Vienna
  • 4. University of Leipzig
  • 5. University of London

Description

KITAB Text Reuse Data

 

KITAB’s text reuse data is generated by running passim on the OpenITI corpus (DOI: 10.5281/zenodo.3082463). Each version is the output of a separate run. To prepare the corpus for a passim run, we chunk texts into passages of 300 tokens (~words) in length. Also, we normalize texts and remove all non-Arabic characters. The chunks, called milestones, are identified by unique ids. This dataset represents the reuse cases that have been identified among milestones. 

The dataset contains folders for each book. Each folder includes alignment files between that book and all other books with which passim has found instances of reuse. The reuse cases between a pair of books are represented as a list of records. Each record is an alignment that shows a pair of matched passages between two books together with statistics, such as the algorithm score, and contextual information, such as the start and end positions of aligned passages so that one can find those passages in the books. A description of the alignment fields is given in the release notes.

For each dataset, we generate statistical data on the alignments between the book pairs. The data is published in an application that facilitates search, filtering, and visualizations. The link to the corresponding application is given in the release notes.

KITAB is funded by the European Research Council under the European Union’s Horizon 2020 research and innovation programme, awarded to the KITAB project (Grant Agreement No. 772989, PI Sarah Bowen Savant), hosted at Aga Khan University, London. In addition, it has received funding from the Qatar National Library to aid in the adaptation of the passim algorithm for Arabic.

Note on Release Numbering: Version 2020.1.1—where 2020 is the year of the release, the first dotted number—.1—is the ordinal release number in 2020, and the second dotted number—.1—is the overall release number. The first dotted number will reset every year, while the second one will continue on increasing.

Note: The very first release of the KITAB text reuse data (2019.1.1) is published here as it was too big to publish on Zenodo. To receive more information on the complete datasets please contact us via kitab-project@outlook.com (or other team members). 

Future releases may include part of the generated data if the size of whole data is too big to publish on Zenodo. However, the data is open access for anyone to use. We provide the detailed information on the datasets in the corresponding release notes.

Files

KITAB-reuse-data_2020-1-2.zip

Files (48.1 GB)

Name Size
md5:bae87c9d84930f7ba6c5619b29fef469
48.1 GB Preview Download

Additional details

Funding

European Commission
KITAB - Exploring Cultural Memory in the Pre-Modern Islamic World (700–1500): Knowledge, Information Technology, and the Arabic Book 772989