OpenITI MAKHZAN
Parkes Allen, Jonathan; Mullan, John; Nigst, Lorenz; Barber, Mathew; Shahid Khan, Taimoor; Seydi, Masoumeh; Chen, Danlu; Weng, Yufei; Vogler, Nikolai; Murel, Jacob; Eshera, Osama; Berg-Kirkpatrick, Taylor; Smith, David; Bowen Savant, Sarah; Thomas Miller
- Publisher
- Zenodo
- Language
- en
Abstract
OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript Data The Open Islamicate Texts Initiative (OpenITI) was founded in 2017 with the goal of building the digital infrastructure for the study of the premodern Islamic world. The underperformance of automatic transcription via Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) for Arabic-script languages long stymied the production of large-scale corpora in Islamicate languages, such as Persian, Arabic, Ottoman Turkish, and Urdu. The OpenITI MAKHZAN dataset is a large aggregation of Arabic-script ground truth and evaluation data drawn from a wide variety of Persian, Arabic, Ottoman Turkish, and Urdu print and handwritten (manuscript) documents. Note on Release Numbering: Version 2025.1.2—where 2025 is the year of the release, the first dotted number—.1—is the ordinal release number in 2025, and the second dotted number—.2—is the overall release number; the first dotted number will reset every year, while the second one will continue on increasing.
Full text
OpenITI Makhzan, ver. 2025.1.1 —Release Notes— The current publication is the first release of OCR/HTR output of selected print and manuscript pages by the Open Islamicate Texts Initiative (OpenITI) using eScriptorium. The release includes: ● data: main data folder (OpenITI-Makhzan_Data_2025-1-1.zip) containing zip files for each document (i.e., book or manuscript). ○ Each ZIP file is named after the document id assigned by eScriptorium, and is preserved in the “Doc ID” column in the metadata file. ○ Each ZIP file contains image and XML files of the document’s pages, whose filenames include an eScriptorium-assigned unique number preserved in the “Doc Part ID” column in the metadata file (i.e., file names follow “Doc ID_Doc Part ID” pattern). ○ Each ZIP file (document folder) also includes a mets.xml file that holds the information of the files of the current document. ○ ZIP files can be directly imported into eScriptorium as documents.1 ● metadata: metadata file (OpenITI-Makhzan_Metadata_2025-1-1.tsv) ○ Each row in the metadata file corresponds to a single page (document part) in the associated ZIP file of the document. You can find the column descriptions for the metadata file, detailed in this document. ○ All pages belonging to the same document share the same document id (“Doc ID”). ○ Each page has its own unique “Doc Part ID,” contained in the file name of that page’s image/XML pair; ● release_notes: OpenITI-Makhzan_ReleaseNotes_2025-1-1.pdf. 1 In case the import fails, try to unzip and re-zip the file and then re-import the new zip file.
Metadata Column Descriptions ● Doc ID: Primary key of the eScriptorium document from which the page or document part has been harvested. ● Doc Part ID: Primary key of the individual page or part within the corresponding document. ● Link to Doc Part: Direct link to the document part/page in eScriptorium. ● Transcription Layer Name: Name of the transcription layer within the document part. ● Language: Language of the text on the page. ● Lines: Number of text lines on the page. ● Script: Script style of the text (e.g., naskh, nastaliq). ● URI: OpenITI URI of the text (given if the text is part of the OpenITI corpus). ● Title of Volume: Title of the volume (manuscript or print) containing the text; one volume may contain multiple texts. ● Title of Text: Title of the individual text. ● Place of Origin: Place of origin or production of the volume (manuscript or print). ● Call No: Shelf mark or call number of the manuscript. ● Repository: Holding institution of the manuscript. ● Place of Repository: Geographic location of the holding institution. ● Public Link: Static URL to the source manuscript (if available online). ● Author/Compiler of Volume: Name of the author, editor, or compiler of the volume. ● Author/Compiler Dates: Birth and death dates of the author/compiler of the volume. ● Author of Text: Name of the author of the text. ● Author Date: Birth and death dates of the text’s author. ● Volume Date: Date of the volume’s compilation or production. ● Type: Page type (e.g., publication [kitābat], manuscript). ● Form: Literary or visual form of the text (e.g., prose, poetry, title page). ● Genre: Literary or intellectual genre of the text (e.g., history, biography, philosophy). ● Category: Vetting status of the transcription. ● Segmentation Complete: TRUE/FALSE indicating whether all lines on the page have been segmented. ● Transcription Complete: TRUE/FALSE indicating whether the transcription of the page is complete. ● Transcription: Name of the person who performed the transcription. ● Vetting: Name of the person who vetted the transcription.
● Notes: Miscellaneous notes (e.g., tight page spacing, columnar layout, marginalia).