Evaluating morphosyntax in an evolutionary perspective
Abstract
This presentation informed experts on computational phylogenetics as applied to historical linguistics from different fields about PhylProGramm's objectives and methods.
Full text
Evaluating morphosyntax in an evolutionary perspective Elena Anagnostopoulou C-LESTE Tuebingen, June 15-16 2024
Common traits of languages can have four different sources: •Common ancestry •Geographic proximity/contact •Universal trends •Chance similarities
The PhylProGramm project aims to: ●evaluate morphosyntax with respect to these dimensions ● determine the potential of individual features for tracking language evolution ●test the hypothesis that there are correspondences between sizes of Chomskyan parameters (macro-parameters, meso-parameters, micro-parameters) and degrees of temporal stability (Roberts 2019).
A study of morphosyntax via phylogenies ● With the aid of a combination of evolutionary biology-inspired and linguistic methods, we aim to: ●disentangle the mixed signal of morphosyntactic features in order to ○ identify and explain their resistance and propensity to change and borrowing ○ at the scales of macro-variation, meso-variation and micro-variation.
Context: debates & controversies Morphosyntactic data have not been scrutinized in ways analogous to cognate data (Swadesh lists).
What kind of morphosyntactic data are adequate for phylogenies? ● Typological? ● Parametric? ●Other, e.g. encoding cognacy? ● Inadequate?
Are stability and diffusibility intrinsic and constant properties of features? Nichols 1992, 1998, 2003 Inherit Borrow Substratum Select Basic vocabulary High Low ? n/a Ergativity Low Low High? Low Gender Not high Low ? Nil? Inclusive/Exclusive High Appreciable High? Low SOV High High High? High? Verb-initial Low Low High? Low
Are morphosyntactic data temporally more or less stable than cognate data? Morphosyntax generally faster than lexicon (Greenhill et al. 2017) Morphosyntax may be the key for deep relationships (Nichols 1992, Matsumae et al. 2021, Greenhill 2021, Santos et al. 2020)
Is there arbitrary variation across families? ●Are rates of change, evolutionary dependencies and transition probabilities constant and predictable or do they vary arbitrarily across branches, families and areas? ● On rates of change, compare the consensus stability rankings in Dediu & Cysouw (2013) based on global samples to Greenhill et al.’s results for Austronesian. ● On lineage-specific trends in word order Universals, see Dunn et al. (2011) and related work on Indo-European, Austronesian, Bantu and Uto-Aztecan.
The e-atlas of Cappadocian Dialects Dimitra Melissaropoulou and Stavros Bompolas http://cappadocian.upatras.gr/en/n ode/10 Melissaropoulou et al. (2022)
Detailed documentation of phonology, morphology, syntax, lexicon, and their geographic distribution word + -ja stem-allomorph + -ja word + -ja word + -ja stem + -ja stem + -i stem-allomorph + i 19 varieties, 279 phenomena classified as Contact-Induced (CI) vs. Language Internal (LI). For each feature: detailed comments and references (1st column), instantiations (2nd column), examples (3rd column)
-Dialectometric comparisons of phonology, morphology, syntax. -Dialectometric estimation of degrees of association between CI and LI features, and degrees of correlation between geography and dialect areas . -Bompolas & Melissaropoulou 2023, forthcoming; Bompolas 2023
Planned and current work ● We will put the existing synchronic data in a diachronic perspective, exploring the interaction between LI and CI phenomena with phylogenetic (and perhaps also phylogeographic) techniques. ●A reference tree for Inner Asia Minor Greek (and later all dialects) based on cognates is under construction. ● We will expand this line of work to the other dialect groups via morpho-syntactic (and also phonological) data collection and curation at a feasible level of granularity, taking into account geographical, typological and sociolinguistic parameters (interactions between typological compatibility, and intensity of contact, e.g. comparison of Italiot Greek varieties and Cappadocian varieties spoken in nearly identical sociolinguistic settings). ● As a side project: Do the trees provide evidence for a Medieval Koine ancestor or was Hellenistic Greek the last Koine? ● Group members involved: Dimitra Melissaropoulou, Stavros Bompolas, Natalia Chousou-Polydouri, Giorgos Kostopoulos (currently; later also DM, EA, 1 PhD student)
Meso-variation: Indo-European ●46-IE-Language Sample (complete) ○ Sample: 46 languages for which cognate data are available ●62-IE-Language Sample ○62 languages (superset of 46) Wals-based characters Multivalued and binarized Modified and expanded Data collected from grammars, PhD dissertations, articles. Group members involved: Maria Margarita Makri, Ruby Sleeman, Elena Anagnostopoulou, Dimitris Michelioudakis, Stergios Chatzikyriakidis
Cognate vs. Morpho-syntactic tree Cognate tree Morpho-syntactic tree
Deviations from received wisdom Geography/Contact Alb, Rm, Grk, Blg, Mac form a group. Sprachbund. Homoplasies English should be attached to German/Dutch, not Danish/Swedish. Old languages Sanskr/Got/OE/Lat/CIG are grouped together. Interestingly, Greek dialects correctly form a group with Standard Modern Greek. Computational methods to detect the sources for these deviations Some methods were developed and applied in the ModelGloss project. Group members involved: Manolis Ladoukakis, Pavlos Pavlidis, Christos Zioutis. For example:
Bipartitions method 1 1 0 0 Bipartition labels b1 Bipartition labels b2 Character states 10 0 r2 1 B C DA 23 Markers for groupings of interest i.e. Sprachbunds, Celtic, Romance, Italo-Celtic, English-Mainland Scandinavian, Old systems B C DA 1 1 1 0 Language B C D A r1
Macro-variation ● Detailed investigation and comparison of the dimensions Inherit/ Borrow/ Select in morphosyntactic features of old, large and known families: ○ Indo-European ○ Austronesian ○ Niger-Kongo ● Currently: scrutinizing and processing existing cross-linguistic databases (WALS, Autotyp, Grambank, SSWL, TypLinkInd (TLI) (Graff et al., in prep), DiACL, ModelGloss) in terms of informativity of features, mutual compatibility and combinability, data coverage, interdependencies of features and missing values. ● Group members involved: Maria Margarita Makri, Ruby Sleeman, Natalia Chousou-Polydouri, Stergios Chatzikyriakidis and Elena Anagnostopoulou
Development of a simulator Simulator that includes processes allowing us to test the linguistic data we are interested in. For example: ● horizontal transfer ● multistate features ● recombination blocks (some features may have the same history, while others different history) ○ for example, some features horizontal transfer, others vertical transfer ● ancestral languages with polymorphism ● test whether multistate or binarized features result in more accurate trees ● infer the amount of horizontal transfer to obtain “Sprachbunds” ● understand the effect of hierarchies in data Group members involved: Pavlos Pavlidis, Manolis Ladoukakis, Christos Zioutis, Stergios Chatzikyriakidis
GenN and agreeing head marking •EMPIRICAL FINDING: Agreeing head-marking is permitted in N-GEN orders as well. • Thus far, we have identified 47 such languages out of a sample of approximately 155 with agreeing head-marking (from the Noun Possession Classes database; Chousou-Polydouri, Inman et al., 2023; Out of Asia project, UZH ). •EMPIRICAL GENERALIZATION: In agreeing head-marking languages, Gen/Poss is treated (also order-wise) like the S argument. • This result was based on our findings from the database we constructed. • In particular, two findings were important: •N-GEN orders with agreeing head-marking typically align with VS orders; NB: including languages where transitive constructions follow SVO order but intransitive ones adopt VS structure. • Head-agreement possessor marking always aligns with the marking of S in the verbal domain---and those cases in which it aligns with P are only apparent counter-examples (e.g. active-stative alignment).
Conclusion ⇒ Agreeing head-marking is a good predictor of the relative ordering of Gen and N, once you know the constituent order and alignment in clauses • How do we test the stability of ‘agreeing head marking’? Are the methods (i) applicable, (ii) suitable? • Also, how much is “agreeing head marking” represented in the global samples? And how well does it correlate with “Gen N” in these samples?
Parameter sizes, pervasiveness and stability A research plan
Roberts 2019: a model of variation ●Roberts (2019; synthesizing work of the ReCoS project and its team) proposes a model of variation aiming to: ○Resolve the empirical and conceptual problems associated with classic Parameter Theory (Chomsky 1981; Chomsky & Lasnik 1993). ○Reconcile three conceptions of Parameters that may capture the full space of variation we observe: ■Macro-Parameters (Baker 1996, 2008) ■Micro-Parameters (Kayne 2000) ■The “Borer-Chomsky conjecture”, namely that language variation derives from properties in the functional inventory of languages (thus reflecting variation in the lexicon)
Microand Macro-parameters ●In a nutshell, Roberts proposes: ○Language variation derives from differences in the feature specification of functional heads (Kayne’s microparameters & Borer’s conjecture). ○Macroparameters (in the sense of Baker) are derivative: they arise when all functional heads in a language are ‘harmonic’ w.r.t. a designated feature. ○Macroparameters are aggregates of microparameters acting together Microparameters Macroparameters Affect a small number of functional heads Affect all relevant functional heads Not pervasive in Primary Linguistic Data Salient in Primary Linguistic Data Diachronically unstable Diachronically stable
A four-way taxonomy ●eThe existence of ○a) intermediate cases (smaller than macro-parameters and bigger than micro-parameters) ○b) highly unsystematic variation associated with particular lexical items leads to a four-way typology of parameters: ■Macro-Parameters: properties maximally characterizing the functional inventory of a language (cf. Cross-categorial harmony in typology) ■Meso-Parameters: properties characterizing large natural classes of heads, e.g. nominal or clausal. ■Micro-Parameters: properties characterizing small, lexically definable classes of functional heads. ■Nano-Parameters: unsystematically expressed on individual lexical items/ irregularities.
Parameter Hierarchies ●They restrict the space of possible grammars and have the following form (word order example): a. Is F present in the system? (Y/N) If N: rigidly, harmonically head initial language b. If Y: is F generalized to all (relevant) heads? Y: rigidly, harmonically head-final language c. If N: is F restricted to some subset of heads, e.g. [+V]? d. If N, further restrictions on the distribution of F tThe presence of a feature F in different parts of speech and/or different constructions can be viewed as feature pervasiveness or feature frequency in different theoretical frameworks.
Metric(s) of pervasiveness/parameter size Plan: ●Quantify Roberts’ insights into a ●Continuous character that reflects the parameter size ○ Head-directionality ○ null subjects ○ case and alignment ●Compare the results to alternative approaches, e.g. functional approaches Group members involved: Natalia Chousou-Polydouri, Dimitris Michelioudakis, Pavlos Pavlidis, Manolis Ladoukakis, Nikos Angelopoulos, Elena Anagnostopoulou
Evolution of metrics on phylogenies We expect clear optima when something is a macro-parameter and/or something is pervasive ● Fitting of a Ornstein-Uhlenbeck model ● Comparison with Brownian Μotion
Fitting of evolutionary models ● Known phylogenies ○ Indo-European ○ Austronesian ○ Niger-Congo ● Models of evolution ○ Brownian motion ○ Ornstein-Uhlenbeck Focus on 2-3 features from Roberts. Collection of fine grained data. ○ Head-directionality ○ null subjects ○ case and alignment Blanquart F, Wymant C, Cornelissen M, Gall A, Bakker M, et al. (2017)