scieee AI-readable full text Open interactive document viewer

The Greek New Testament Text-Fabric Conversion Project (Navigating the Challanges in Creating a Text-Fabric Dataset)

Jurg, Tony

Abstract

This presentation discusses the methods used in the 'Greek New Testament Conversion Project,' which aimed to create a Text-Fabric dataset representing the Greek New Testament. It will detail the considerations made during the preparatory phase of the project, including defining the requirements and selecting the most appropriate data source, along with its provenance. The presentation further delves into the various implementation attempts using different conversion methods integral to the conversion project, evaluating their suitability for our project and in a broader context. Some technical issues encountered will be discussed, along with how these challenges were resolved. One of the primary requirements for the Text-Fabric implementation was the capability to display syntax trees. This led to the creation of a unique feature that allows for the display of these trees in either a 'word-group' or a 'syntactic' view, selectable through a simple command. The implementation of this feature will be discussed at a high level, showcasing the flexibility and extensibility of Text-Fabric. Additionally, attention will be given to the importance of creating documentation and how some tools created as a spin-off of this project can assist in this process.

Full text

The Greek New Testament Conversion Project Session 30-24 (joint session with Eep Talstra Centre for Bible and Computer and Hellenistic Greek Language and Linguistics) ●Willem Th. van Peursen, Vrije Universiteit Amsterdam and Oliver Glanz, Andrews University: The Greek TF Conversion Project ●Tony Jurg, Vrije Universiteit Amsterdam: Navigating the Challenges in Creating Text Fabric Datasets (with Focus on the NT Greek Corpus) ●Saulo de Oliveira Cantanhêde, Andrews University: Coding the Greek New Testament in an Open Source Environment SBL International Annual Meeting - July 30, 2024 Navigating the Challenges in Creating Text - Fabric Datasets Tony Jurg, ETCBC/VU Amsterdam This presentation Share some highlights: •How this projects was executed •Key decision points •Learnings that can benefit other projects 'GNT TF Conversion Project’ Objectives Create a High-Quality Text-Fabric Dataset: a Text-Fabric dataset of the Greek New Testament that matches the usability and quality standards of the BHSA. Comprehensive Documentation: quality end-user focused documentation demonstrating and unlocking the dataset potential. Publications and Knowledge Sharing: •Conversion Project Insights: An article accounting for the conversion process, process from both technical and organizational perspectives; highlighting successes, challenges, and lessons learned. •Dataset Overview: An article with comprehensive description of the resulting database, enabling prospective users to grasp its core features and potential applications. •Educational Applications: An article on how this Text-Fabric dataset can be utilized in educational settings, particularly for teaching Biblical Greek. •Syntactical Analysis: An article on leveraging this Text-Fabric dataset for syntactical analysis in Biblical Greek studies, demonstrating its benefits for linguistic research. Preliminary Considerations First Question: Why pursuing this? •Does it unlock new types of data previously unavailable? •Does it enable or enhance original research options? •Does it expand analytical possibilities? •Does it make the data more accessible or easier to use? •Will it increase your audience? •Will it improve transparency, reproducibility and accountability •Any other benefits of the data conversion? For our project: The primary benefit (or requirement) was to obtain a quality query and display function for syntax trees in a Greek New Testament Text-Fabric dataset that is open access and non-restrictive licensed. Preliminary Considerations Second Question: Is it worth the effort? •Are the necessary resources such as funding, access to data, and time available for the conversion? •Does the anticipated benefit justify the investment of resources? •Can the project be sustained over time with the available resources? •Are there any legal, ethical, or regulatory requirements that need to be resolved? For our project: The project deliverable (a new function for syntactic queries on constituency trees) justifying the investment. Additionally, the project's execution became a research subject for an internship. Data Source Assessment Examine key aspects of prospect data sources: •Provenance •Stability •Quality •Sustainability •Copyright status For our project we used ‘Nestle 1904’ because: •Public domain •Preexisting XML Treebank with constituency grammar •Word level compatibility with Bible Online Learner •Nestle 1904 de facto standard for linguistic research XML input data stability Typical XML element structure: <tag attribute1="value1“ attribute2="value2">text</tag> XML Input Data Stability impact: •Attribute changes: Orthographic/morphological changes have minimal impact; syntactical changes can significantly alter presentation and query results. •Tag name changes: Can cause compatibility issues, requiring updates to the conversion tool. •Tree presentation changes: Affect element arrangement and relationships, potentially altering data interpretation. Database organisation: node types Important: Text-Fabric features are attached to nodes Design questions: •Which input data elements should translate to nodes? •Which node types need to be added For our project: •Dropped some tags: e.g. the <p> and the <milestone> tags from the LowFat Tree data. •Added corpus organizational node types that were not explicit in the XML data like ‘book’, ‘chapter’, and ‘chapter’. • Added syntactic node types like ‘phrase’ or ‘clause’ mirroring the ‘wg’ nodes. Method 2: Intermediate relational database This method uses TF's walker function, requiring preprocessing of XML source data stored in a MariaDB database. Each tag-type has its own table, with attributes as fields. The director function extracts and feeds the data to the walker, decoupling the XML structure from the walker’s input format. Method 3: Intermediate Flat Data Tables A script collects XML data into a pandas DataFrame for each source file, with each <w> tag becoming a row and its attributes forming the columns. This converts the structured XML into flat tables with redundancy. The DataFrames are saved as binary pickle files (.pkl). In the second stage, the pickle data is converted into TF format using the tf.convert.walker module, guided by the director function. Coding environment Software Stack Development Environment Coding environment: OS impact Potential GitHub problems due to OS interactions: Warning: the following paths have collided (e.g., case-sensitive paths on a case-insensitive filesystem), and only one from the same colliding group is in the working tree. File paths separators back slashes forward slashes Home directory %USERPROFILE%. tilde (~) Case handling case-insensitive case-insensitive handling of next line \r\n \n Open TF environment Open to all sides: •Data sources (in our dataset: optional add-on features) •Additional functionality (in our dataset: view type commands) •Flexible post processing of resulting data • … etc. Unlocking potentials Problem: access to data does not give understanding Mitigation: •Clear and consistent end-user documentation •Dense Hyperlinking •Speeding up learning curve by providing examples A well-documented dataset with comprehensive metadata allows researchers to easily understand its content, potential and value Feature documentation End-user documentation maintenance tonyjurg/Doc4TF Observation: Syncing feature documentation with development is labourintensive and error-prone due to various types of updates, such as: •Adding or deleting features or changing feature names •changes in associated node types •Changes in value domains FAIR principles https://www.go-fair.org/fair-principles Featured Image: FAIR data principles by SangyaPundir under CC BY-SA 4.0 license Published location https://centerblc.github.io/N1904/ centerblc/N1904