scieee AI-readable full text Open interactive document viewer

Visualising Overlooked Syntactic Alternatives in the Greek New Testament

Jurg, Tony

Abstract

Paper presented at the IOSOT 2025 conference in the section Digital Humanities and Computational Approaches to the Bible. This paper explores how overlooked syntactic alternatives in the Greek New Testament can be visualized by extending the N1904-TF dataset. By integrating Morpheus-derived morphological analyses through a custom pipeline, multiple possible parses for each word become visible, showing how ambiguity at the morphological level can shift phrase functions and reveal alternative readings often hidden in disambiguated datasets. While this method enriches grammatical exploration, it also exposes inherent tensions: N1904-TF assigns one disambiguated morphosyntactic tag per token, whereas the Morpheus data records all legitimate possible renderings, including those resulting from underspecification. The principles of contextual adjudication versus recording underspecification therefore result in incompatibility when performing comparisons. The proposed way forward is to codify underspecification (distinguishing all possible values from the chosen contextual reading) and to expand the training data with corpora like the Septuagint, paving the way for more precise and transparent computational study of the New Testament

Full text

T. Jurg – IOSOT 2025 - Visualising Overlooked Syntactic Alternatives in the GNT page 1 Visualising Overlooked Syntactic Alternatives in the Greek New Testament 25th Congress of the International Organization for the Study of the Old Testament. Berlin, Humboldt-Universität zu Berlin, Faculty of Theology, 11–15 August 2025. Session: Digital Humanities and Computational Approaches to the Bible 2 (42-108, Thursday, 08/14/2025). Paper presented by Tony Jurg (Vrije Universiteit Amsterdam / ETCBC).1 Sheet 1: Cov er Sheet I am honoured to present at this year’s IOSOT conference the research I have been developing over the past years. As the title suggests, this paper addresses the visualization of syntactic ambiguity in the Greek New Testament. Such a topic is wide-ranging and touches on many dimensions, but within the time available I will necessarily focus on selected aspects. Over the next twenty minutes, I will begin by situating the project in its broader context, then highlight several key results, and finally conclude with some critical observations and forwardlooking proposals for further research. Sheet 2: Th e N1904-TF D ataset 1 ORCID: https://orcid.org/0000-0002-0343-1346. Academia: https://vu-nl.academia.edu/TonyJurg. T. Jurg – IOSOT 2025 - Visualising Overlooked Syntactic Alternatives in the GNT page 2 At last year’s SBL International Meeting, we introduced a digitally annotated edition of the Greek New Testament. Developed jointly by the Eep Talstra Centre for Bible and Computer (ETCBC)2 at the VU Amsterdam and the Center of Biblical Languages and Computing (CBLC) at Andrews University (Berrien Springs, Michigan),3 this project transformed the Nestle 1904 (7th ed.) into a constituency-parsed, syntactically annotated dataset.4 While certainly not the first digital annotated edition of the GNT, its significance lies in bringing the New Testament into the same Text-Fabric ecosystem,5 which was already available for the Old Testament (BHSA).6 In that sense, the New Testament has now caught up. Our dataset was deliberately designed to mirror the BHSA in both taxonomic structure and user experience. With more than fifty features, the dataset offers a comparable depth of annotation, all accessible through Text-Fabric’s versatile API and accompanied by graphical syntax-tree visualizations. The N1904-TF dataset is fully documented and released under a permissive open-source license.7 Sheet 3: Tex t-Fabric Graph Structure (N190 4-TF) One of Text-Fabric’s core strengths as a research ecosystem for textual analysis lies in its graphbased architecture. This design cleanly separates the surface text from its annotations, ensuring both flexibility and extensibility. Each linguistic unit (whether a word, phrase, clause, or sentence) is represented as a node, identified by a unique integer within the dataset’s namespace. At the lowest level are the slots, which typically correspond to surface-level elements. These slots are the smallest addressable units in the dataset and, in NLP terminology, are often referred to as tokens. Higher-level structures, such as phrases, clauses, or sentences, are modelled as additional nodes linked to these slots, enabling a hierarchical organisation of annotations. All descriptive information is stored as features, capturing properties like surface text, lemma, part of speech, 2 https://etcbc.nl/. 3 https://www.andrews.edu/. 4 The source data for the conversion project was the macula-greek version of Eberhard Nestle’s 1904 Greek New Testament, published by the British and Foreign Bible Society. The digital source is available at GitHub: Clear.Bible, “Clear-Bible/macula-greek,” accessed June 25, 2025, https://github.com/Clear-Bible/maculagreek/tree/main/Nestle1904/, (licensed under CC BY 4.0). The dataset is based on the print edition: Eberhard Nestle, Η Καινή Διαθήκη: Novum Testamentum Graece (New York: Fleming H. Revell Company, 1904), with additional text-critical markers and associated texts drawn from the seventh reprint (1913). 5 Text-Fabric is a Python package created by Dirk Roorda. It offers a research ecosystem for analysing and manipulating large textual datasets, especially suited for ancient languages and biblical texts. The software is available at https://github.com/annotation/text-fabric/. 6 https://etcbc.github.io/bhsa/. 7 https://centerblc.github.io/N1904/. T. Jurg – IOSOT 2025 - Visualising Overlooked Syntactic Alternatives in the GNT page 3 morphological tags, and syntactic functions. Conceptually, a feature is a key–value mapping that links a node ID to its property value.8 The accompanying image, which visualises the graph representation of the first clause of John 1:1, demonstrates how nodes and features interconnect. Sheet 4: Tex t-Fabric Annota tion Layers (N 1904-TF) This graph-based architecture also enables the seamless integration of annotations. In the initial release of the base dataset, N1904-TF, fifty features were included, spanning multiple linguistic levels: from orthography and lexis to morphology, syntax, and even higher-level semantics such as semantic role labelling. Crucially, the architecture also supports the integration of additional datasets consisting solely of features, which attach new annotations to the existing nodes via their unique identifiers. For example, we released a supplementary dataset alongside N1904-TF that builds directly on this principle: it adds new annotation layers without altering the underlying text. Its features added further granular annotations, such as declension class, crasis, or discourse details like Synoptic Gospel parallels.9 All annotations remain uniformly accessible through the Text-Fabric API, allowing scholars to construct complex queries across arbitrary sections of the data by combining features from both the base dataset and any load-on-demand extensions. 8 Dirk Roorda, ”Text-Fabric Data Model,” accessed June 27, 2025, https://annotation.github.io/text-fabric /tf/about/datamodel.html. See also section “Logic: annotated graphs” on the data model of Text Fabric in Dirk Roorda, “Text-fabric: handling Biblical data with IKEA logistics,” HIPHIL Novum vol 5 issue 2 (2019): 126–135, 128, doi.org/10.7146/hn.v5i2.142740. 9 https://centerblc.github.io/N1904/additions/index.html. T. Jurg – IOSOT 2025 - Visualising Overlooked Syntactic Alternatives in the GNT page 4 Sheet 5: Tex t-Fabric Syntax Tree (N1904-TF) Our Text-Fabric dataset was originally derived from the Macula XML corpus, which is organised around word-groups. To bring it closer to the BHSA model, we added a parallel syntax-oriented layer that labels phrases and constituents in BHSA style (e.g., Subject, Predicate). In Text-Fabric, syntax trees are visualised as nested boxes, each box representing a linguistic unit with its properties listed inside. This sheet shows the syntax tree for John 1:1 in a BHSA like style. Rather than enforcing a single linguistic framework, we encoded both representations side by side and provided tools that allow users to toggle between the word-group view and the syntax-tree view. This dual presentation was made possible by duplicating the relevant data structures and introducing shadow nodes to capture the alternative perspective.10 Having both views available raised an important question: could such dual encoding also shed light on cases of alternative punctuation, such as the long-standing debate over John 1:3–4?11 That dispute stems from a phrase-attachment ambiguity: the grammar permits multiple readings, but the text itself does not resolve them. This observation prompted a search for comparable configurations elsewhere in the New Testament and for ways to make such alternatives visible. In principle, multiple syntactic interpretations may be both grammatically valid and attested in the corpus, yet they often remain hidden from the reader. To address this, I experimented with encoding multiple alternative syntax trees, including variations at the phrase, clause, and sentence boundaries, stored within a single “syntax forest.” A small demonstration dataset with companion code confirmed that Text-Fabric can accommodate such parallel encodings, enabling queries to retrieve matching patterns across multiple trees using the standard syntax.12 The natural next step was to scale this up to the full dataset. For that purpose, I trained a Probabilistic Context-Free Grammar (PCFG) on the full N1904-TF corpus in order to discover alternative syntactic parses. For this I modified NLTK’s Viterbi parser to return probability-ranked sets of parses. As is typical for constituency parsing, the process proved computationally slow. 10 https://centerblc.github.io/N1904/viewtypes.html. 11 For details on reception history and the possible syntactic interpretations of these verses, refer to John F. McHugh's 'A Critical and Exegetical Commentary on John 1–4,' edited by Graham N. Stanton and G. I. Davies, in the International Critical Commentary series (London; New York: T&T Clark, 2009), 14. For an analysis of the theological implications of the larger text, refer to Martinus C. de Boer, 'The Original Prologue to the Gospel of John,' in New Testament Studies 66 (2009), 448-467. 12 https://github.com/tonyjurg/MPTF. T. Jurg – IOSOT 2025 - Visualising Overlooked Syntactic Alternatives in the GNT page 5 However, a more significant issue emerged. Review of the output revealed a fundamental limitation: a bare “forest” of alternative trees could illustrate the range of possible outcomes but offered little insight into the underlying reasons. This limitation was already present when working with the disambiguated morphological tags, but it became even more acute once morphological ambiguity was included. It came to a point where the number of parallel syntax trees literally exploded, until it was impossible to see the forest for the trees. Sheet 6: Morp heus Morphologic al Analys er This prompted a shift in perspective: rather than beginning with the effects, an ever-expanding forest of syntax trees, why not return to the underlying cause? What if the first step were to focus on morphological ambiguity itself, integrating it into Text-Fabric, and then developing diagnostic features to detect, visualize, and even quantify its impact on phrase and clause structure? Such an approach promised far clearer insight than attempting to sift through an overwhelming array of competing parses of full sentences. This refocusing placed morphology at the centre of the project. To generate all possible morphological parses for each wordform, I selected Morpheus, a tool noted for its broad dialectical coverage and established reputation.13 Its development reflects decades of work by Gregory Crane and the Perseus Project in advancing accessible computational resources for classical languages. The parser uses a rule-based system based upon extensive lexical tables, encompassing “40,000 stems, 13,000 endings, and 2,500 irregular forms.”14 13 https://github.com/PerseusDL/morpheus/blob/master/doc/morpheus.html. 14 Gregory Crane, “Generating and Parsing Classical Greek,” Literary and Linguistic Computing, Vol. 6, No. 4,1991: 241– 245, 244. doi:10.1093/llc/6.4.243. T. Jurg – IOSOT 2025 - Visualising Overlooked Syntactic Alternatives in the GNT page 6 Sheet 7: Hig h Level Productio n Pipeline To capture all ambiguity data for each word, I designed an NLP pipeline with reproducibility as a central principle.15 The system consisted of three distinct main components: a Jupyter Notebook for generating the Text-Fabric features, the Morpheus analyser for producing morphological parses, and a dedicated gateway that linked the two. The N1904-TF dataset supplied the surface wordforms, which were passed to Morpheus for analysis. Since Morpheus depends on an older software version, I ran it in a virtual machine,16 accessed through a lightweight API. To bridge the gap between Morpheus and Text-Fabric, which rely on very different taxonomies and data structures, I also developed a dedicated Python package called Morphkit.17 Morphkit managed all interactions between the two environments by parsing each analytic block produced by Morpheus, delivered as plain-text with fixed-position tab-separated fields, and converting it into structured Python dictionaries suitable for Text-Fabric feature generation. It also ensured data consistency by aligning and mapping property taxonomies. Sheet 8: Th e Morpheus A ddons Dataset Executing this NLP pipeline produced a Morpheus-enhanced dataset for N1904-TF. For each token in the base dataset, all morphological details from every Morpheus block were stored in dedicated Text-Fabric features. Because Morpheus can return up to 24 analytic blocks for a single 15 https://tonyjurg.github.io/Create_morpheus_TF_dataset/. 16 https://hub.docker.com/r/perseidsproject/morpheus-api. 17 https://tonyjurg.github.io/morphkit/. T. Jurg – IOSOT 2025 - Visualising Overlooked Syntactic Alternatives in the GNT page 7 token, each representing a distinct morphological parse, additional summary features were introduced to group the analyses by lemma. Metadata features were also added to aggregate counts, such as the total number of blocks per token and the number of distinct lemmas. In addition, several experimental diagnostic features were created, one of which will be discussed in more detail later. Once the Morpheus output was accessible within Text-Fabric, it became necessary to design a method for comparing each morphological parse with the base N1904-TF annotations. To enable this integration, a mapping step was introduced. In N1904-TF, each token is assigned a Robinsonstyle morphological tag, adapted by Ulrik Sandborg-Petersen and commonly referred to by us as the SP tag.18 From the grammatical fields in each Morpheus analysis block, the Morphkit derived one or more SP tags, retaining underspecified values where applicable (e.g., multiple possible genders). This produced a compact and directly comparable encoding of core grammatical properties. Ambiguities became immediately apparent once users enabled the summary features displaying both lemmas and SP tags. Sheet 9: Dat aset Structure and Taxonomy The final Morpheus-enhanced dataset comprises just over 1,000 features, with a total size of about 800 MB. To make it practical for research, we divided it into two parts.19 The lighter part contains metadata, summary features, and analytic features; these are derivative layers designed for quick exploration and broad analysis. The heavier part holds the complete Morpheus output, preserving all analytic blocks and enabling programmatic access to every detail produced by the analyser. This principled division allows researchers to begin with a lightweight dataset and only load the full data when a detailed examination is required. Including all Morpheus data in Text-Fabric is valuable, but on its own it does not necessarily yield insight. For that reason, I created a set of diagnostic and analytic features. These go beyond simple data storage and provide structured ways to explore ambiguity, identify recurring patterns, and detect potential outliers. Together, these features support both programmatic queries and visual inspection, making the dataset a practical tool for studying morphosyntactic alternatives. 18 https://github.com/biblicalhumanities/Nestle1904/blob/master/morph/parsing.txt. 19 https://tonyjurg.github.io/N1904addons/loading_the_dataset.html. T. Jurg – IOSOT 2025 - Visualising Overlooked Syntactic Alternatives in the GNT page 8 Sheet 10: M atthew 2:13 - Tag Sequence Permutations One feature deserves closer attention, as it connects morphological ambiguity with potential shifts in phrase function. The underlying assumption is that a word’s morphological form largely determines its role within a phrase. For example, when a wordform could be analysed either as genitive or dative, this ambiguity might signal a functional shift between possession and agency. Since both the N1904-TF SP tag and the Morpheus parses are available for every token, a Cartesian product can be computed for all items in a phrase. Even a two-token phrase can already yield twelve permutations, each represented by a distinct tag sequence, as shown on the sheet. Sheet 11: M atthew 2:13 - Tag Sequence t o Phrase Functio n Before this step, a mapping table was already prepared. For every unique sequence of morphological tags attested in the N1904-TF dataset, the entire corpus was scanned to determine the observed distribution of phrase functions. These distributions were then stored in a lookup table, linking each morphological tag sequence to the probability distribution of its corresponding phrase roles. The results of the Cartesian product can then be compared against this lookup table. In the example shown, this yields two distinct attestations: one corresponding to the tag sequence found in N1904-TF for the phrase, and another representing an additional alternative tag sequence. Each attestation is associated with a different probability distribution of phrase functions. This comparison provides the basis for testing how ambiguity might shift phrase function. T. Jurg – IOSOT 2025 - Visualising Overlooked Syntactic Alternatives in the GNT page 9 Sheet 12: M atthew 2:13 - Tex t-Fabric Phrases The results are stored in a dedicated set of Text-Fabric features. The baseline feature ma0_pf_altern encodes the original (inherent) parse and its associated probability distribution.20 The feature ma1_pf_altern contains the details for the first alternative tag sequence, and further alternatives can be added in the same way. When these features are displayed, one can quickly examine the potential impact on phrase function if an alternative tag sequence were adopted. In the example, the alternative sequence suggests a possible functional shift from object to subject. Of course, such a shift is not necessarily likely. It depends entirely on context, such as the presence of another phrase in the clause pointing the other way. Another important caveat is that these probability distributions are derived solely from the syntactic parsing of the Greek New Testament, which, in NLP terms, is a very small corpus. As a result, many valid tag sequences are unattested, and the distributions that are available often rest on a limited number of observations. Sheet 13: Eval uation of the M orpheus Data set Now let us evaluate the resulting dataset. The Morpheus-enhanced dataset offers a robust foundation for exploring ambiguity across multiple linguistic levels. The pipeline stores all Morpheus analyses for each word in dedicated features, enabling users to trace results from highlevel SP tags down to individual analytic blocks in a transparent and accessible manner. This approach preserves the full detail of the original data while enhancing usability. The SP tags have proven effective, summarising the detailed output of Morpheus into a single, familiar N1904-TF tag. Because both systems use the same SP-tag taxonomy, comparisons between Morpheus output 20 See https://tonyjurg.github.io/N1904addons/features/ma%7Bind%7D_pf_altern.html.