Full text
The Greek New Testament Conversion Project Session 30-24 (joint session with Eep Talstra Centre for Bible and Computer and Hellenistic Greek Language and Linguistics) ●Willem Th. van Peursen, Vrije Universiteit Amsterdam and Oliver Glanz, Andrews University: The Greek TF Conversion Project ●Tony Jurg, Vrije Universiteit Amsterdam: Navigating the Challenges in Creating Text Fabric Datasets (with Focus on the NT Greek Corpus) ●Saulo de Oliveira Cantanhêde, Andrews University: Coding the Greek New Testament in an Open Source Environment SBL International Annual Meeting - July 30, 2024
Navigating the Challenges in Creating Text - Fabric Datasets Tony Jurg, ETCBC/VU Amsterdam
This presentation Share some highlights: •How this projects was executed •Key decision points •Learnings that can benefit other projects
'GNT TF Conversion Project’ Objectives Create a High-Quality Text-Fabric Dataset: a Text-Fabric dataset of the Greek New Testament that matches the usability and quality standards of the BHSA. Comprehensive Documentation: quality end-user focused documentation demonstrating and unlocking the dataset potential. Publications and Knowledge Sharing: •Conversion Project Insights: An article accounting for the conversion process, process from both technical and organizational perspectives; highlighting successes, challenges, and lessons learned. •Dataset Overview: An article with comprehensive description of the resulting database, enabling prospective users to grasp its core features and potential applications. •Educational Applications: An article on how this Text-Fabric dataset can be utilized in educational settings, particularly for teaching Biblical Greek. •Syntactical Analysis: An article on leveraging this Text-Fabric dataset for syntactical analysis in Biblical Greek studies, demonstrating its benefits for linguistic research.
Preliminary Considerations First Question: Why pursuing this? •Does it unlock new types of data previously unavailable? •Does it enable or enhance original research options? •Does it expand analytical possibilities? •Does it make the data more accessible or easier to use? •Will it increase your audience? •Will it improve transparency, reproducibility and accountability •Any other benefits of the data conversion? For our project: The primary benefit (or requirement) was to obtain a quality query and display function for syntax trees in a Greek New Testament Text-Fabric dataset that is open access and non-restrictive licensed.
Preliminary Considerations Second Question: Is it worth the effort? •Are the necessary resources such as funding, access to data, and time available for the conversion? •Does the anticipated benefit justify the investment of resources? •Can the project be sustained over time with the available resources? •Are there any legal, ethical, or regulatory requirements that need to be resolved? For our project: The project deliverable (a new function for syntactic queries on constituency trees) justifying the investment. Additionally, the project's execution became a research subject for an internship.
Data Source Assessment Examine key aspects of prospect data sources: •Provenance •Stability •Quality •Sustainability •Copyright status For our project we used ‘Nestle 1904’ because: •Public domain •Preexisting XML Treebank with constituency grammar •Word level compatibility with Bible Online Learner •Nestle 1904 de facto standard for linguistic research
XML input data stability Typical XML element structure: <tag attribute1="value1“ attribute2="value2">text</tag> XML Input Data Stability impact: •Attribute changes: Orthographic/morphological changes have minimal impact; syntactical changes can significantly alter presentation and query results. •Tag name changes: Can cause compatibility issues, requiring updates to the conversion tool. •Tree presentation changes: Affect element arrangement and relationships, potentially altering data interpretation.
Database organisation: node types Important: Text-Fabric features are attached to nodes Design questions: •Which input data elements should translate to nodes? •Which node types need to be added For our project: •Dropped some tags: e.g. the <p> and the <milestone> tags from the LowFat Tree data. •Added corpus organizational node types that were not explicit in the XML data like ‘book’, ‘chapter’, and ‘chapter’. • Added syntactic node types like ‘phrase’ or ‘clause’ mirroring the ‘wg’ nodes.
Method 2: Intermediate relational database This method uses TF's walker function, requiring preprocessing of XML source data stored in a MariaDB database. Each tag-type has its own table, with attributes as fields. The director function extracts and feeds the data to the walker, decoupling the XML structure from the walker’s input format.
Method 3: Intermediate Flat Data Tables A script collects XML data into a pandas DataFrame for each source file, with each <w> tag becoming a row and its attributes forming the columns. This converts the structured XML into flat tables with redundancy. The DataFrames are saved as binary pickle files (.pkl). In the second stage, the pickle data is converted into TF format using the tf.convert.walker module, guided by the director function.
Coding environment Software Stack Development Environment
Coding environment: OS impact Potential GitHub problems due to OS interactions: Warning: the following paths have collided (e.g., case-sensitive paths on a case-insensitive filesystem), and only one from the same colliding group is in the working tree. File paths separators back slashes forward slashes Home directory %USERPROFILE%. tilde (~) Case handling case-insensitive case-insensitive handling of next line \r\n \n
Open TF environment Open to all sides: •Data sources (in our dataset: optional add-on features) •Additional functionality (in our dataset: view type commands) •Flexible post processing of resulting data • … etc.
Unlocking potentials Problem: access to data does not give understanding Mitigation: •Clear and consistent end-user documentation •Dense Hyperlinking •Speeding up learning curve by providing examples A well-documented dataset with comprehensive metadata allows researchers to easily understand its content, potential and value
Feature documentation
End-user documentation maintenance tonyjurg/Doc4TF Observation: Syncing feature documentation with development is labourintensive and error-prone due to various types of updates, such as: •Adding or deleting features or changing feature names •changes in associated node types •Changes in value domains
FAIR principles https://www.go-fair.org/fair-principles Featured Image: FAIR data principles by SangyaPundir under CC BY-SA 4.0 license
Published location https://centerblc.github.io/N1904/ centerblc/N1904