Language as a Stack of Homeostatic Property-Cluster Kinds: From Phonemes to Constructions Brett Reynolds ∗ Humber Polytechnic & University of Toronto
[email protected] October 14, 2025 Abstract This paper develops two operational diagnostics – projectibility and homeostasis – for deciding when linguistic categories warrant treatment as homeostatic propertycluster (HPC) kinds. projectibility asks whether a category supports reliable out-ofsample inference; homeostasis asks whether identifiable mechanisms plausibly maintain the cluster over time and across instances. I apply these diagnostics to three structural levels. At the phoneme level I use PHOIBLE inventories to show family-wise concentration of inventory sizes and a scaling relation for the front-rounded vowel /y/; at the lexical level I trace diachronic distributional neighbourhoods to show that some lexemes drift while preserving sufficient cohesion for prediction; and at the constructional level I examine let alone to show that a small bundle of cues transfers across ∗Status note. Conceptual claims are ready for comment; empirical sections contain preliminary analyses pending full verification. Results and figures may change. Feedback on the two diagnostics (projectibility, homeostasis) and the thin/fat/negative failure taxonomy is especially welcome. I got the idea for this paper after reading Ekström et al. (2025), having already read Miller (2021) and chatted with him about his paper and HPCs. I proposed my idea to ChatGPT-5 and had it produce an outline and then a first draft. I used the ChatGPT-5 agent to download the datasets and write and run the python code. Geoff Pullum urged me to “leaven” the text, which was extremely dense. I worked with Clause Opus 4.1 to make the text more accessible to a range of readers. Almost every sentence in this paper was drafted and edited by both of those models. I reviewed, edited, and approved all the material and take full responsibility for the final text and conclusions, which, again, are provisional at this stage. Muhammad Ali Khalidi provided very useful comments on an earlier draft. 1
corpora and degrades predictably under ablation. The contribution is methodological: concrete, reproducible tests that keep kind-claims local and evidence-driven. Where both diagnostics succeed, treating a category as an HPC is empirically warranted; where they fail, a more local or descriptive account is preferable. Keywords: homeostatic property clusters (HPC); linguistic kinds; projectibility; homeostasis; phoneme inventories; PHOIBLE; semantic drift; let alone construction; Universal Dependencies; cross-corpus transfer. 2
Introduction Language presents a familiar problem for cognitive science: the categories we rely on to speak and understand are stable enough to underwrite reliable inference but flexible enough to drift and diversify. Consider how the phoneme /r/ varies across English dialects but remains recognizable, or how the word awful drifted from ‘awe-inspiring’ to ‘terrible’ while keeping its identity. An attractive middle ground for understanding such categories – originating in Boyd’s account of homeostatic property-cluster (HPC) kinds – offers a way to reconcile stability with change (Boyd 1991,1999)1. homeostatic property-cluster kinds emerge when identifiable forces keep characteristic features bundled – not through essence, but through contingent regularities. The category holds together as a family of properties maintained by mechanisms (biophysical, developmental, social) strongly enough to support induction. Think of biological species: robins share characteristic features – size, coloration, song patterns, nesting behaviour – not because of an essential “robin-ness” but because developmental programs, ecological pressures, and reproductive isolation keep these traits correlated generation after generation. Boyd’s insight is that many scientific and social categories work similarly: they’re probabilistic clusters stabilized by mechanisms that keep enough of the cluster together for inductive use. This article proposes a general, testable program: many linguistic categories – phonemes, lexemes, and language-internal constructions – are HPC kinds. The claim is operational, not merely analogical. I test the claim with two simple diagnostics that any proposed linguistic category has to pass. By linguistic category I mean a type maintained by identifiable mechanisms within a population (e.g., the phoneme /r/ in English, a lexeme such as dog, or a named construction such as let alone). For present purposes I treat superordinate labels (phoneme,word) as 1For the first explicit HPC formulation (applied to moral terms), see Boyd (1988: §3.8). For an explicit allowance that social mechanisms can underwrite homeostasis, see Boyd (2000). 3
taxonomic umbrellas; whether they themselves qualify as HPC kinds is an open question requiring different evidence and diagnostics. Here I target language-internal types, where the tests bite. Each case specifies its projection unit (token→token; language→language) and its stabilizers. How long a kind persists, and how widely it extends, are empirical questions about whether stabilizing mechanisms maintain covariance. First, projectibility means that tokens have to support reliable out-of-sample inference about form, meaning, or distribution. Second, homeostasis means that the property cluster must be tied to specific stabilizers with demonstrable covariance. For words, Miller (2021) develops this stance at the level of particular lexemes – dog, run,egregious – rejecting essence-based individuation in favour of mechanism-indexed clusters that are historically delimited and population-relative. On this view, such lexemes earn kind status because mechanisms sustain covarying properties: pronunciation, orthography, meaning, distribution. Miller identifies both cognitive mechanisms (consistent retrieval from mental lexicons) and social mechanisms (community enforcement of usage norms), though he acknowledges the partition depends on theoretical commitments. The lexeme dog, for instance, maintains its identity not through a platonic essence but because spelling conventions, pronunciation norms, semantic associations, and syntactic patterns travel together, stabilized by frequency of use, register and genre licensing, literacy education, and community norms. At the phonological level, recent work by Ekström et al. (2025) argues that phonemes are culturally maintained cognitive tools (Heyes 2018) anchored in articulatory and auditory constraints. Two quantitative signatures exemplify the sort of measurable “homeostasis” HPC requires: family-wise ridgelines of inventory sizes (showing most languages cluster roughly between 20 and 50 segments across unrelated language families; see Figure 2), and a scaling curve in which the probability that a language includes /y/ rises with vowel-inventory size, while /i/ remains common even in small systems. Both patterns are presented with explicit methodological detail and tied to biophysical 4
and efficiency pressures, drawing on data from PHOIBLE, an openly available database of segmental phoneme inventories for over 2,000 languages worldwide (Moran, McCloy & Wright 2019), and a compact checklist of tool criteria summarizes the stabilizers involved (Ekström et al. 2025: Fig. 1 p. 4, Fig. 2 p. 7, Table 1 p. 14). These are precisely the ingredients HPC seeks: projectible distributions and identifiable stabilizing mechanisms. The first case study revisits the phoneme level using the PHOIBLE database introduced above, but focuses on observable patterns rather than re-deriving articulatory physiology. I examine two signatures: how inventory sizes cluster by language family (visualized as ridgelines), and how the probability of finding /y/ in a language increases with the size of its vowel inventory. These patterns survive robustness checks like sampling one language per subfamily. The HPC reading is straightforward: the tight clustering within families demonstrates projectibility (new languages behave like their relatives), while the known mechanisms – quantal stability, dispersion, developmental tuning, and community norms – provide the homeostatic foundation. The second case study turns to words. Building on Miller (2021), I track how a single lexeme’s meaning changes over time (e.g., how egregious drifted from positive to negative) while examining whether its contextual patterns – what words appear near it – remain coherent enough that a model trained on earlier decades can still identify the word in later ones. This demonstrates projectibility empirically. The homeostasis comes from multiple forces: entrenchment through repeated use and norm-guided conventions ensure that even as meanings shift, the word’s spelling, pronunciation, and core distributional patterns stay bundled together at usable levels. The third case study treats an English construction, let alone (Fillmore, Kay & O’Connor 1988), as a language-internal HPC kind. This construction (as in I can’t afford coffee, let alone dinner) signals that the second item is even less likely than the first. Its identifying features – the phrase itself, parallel syntax between the contrasted items, and negative context – work together as a cue bundle that remains stable across different text collections. I 5
test projectibility by training a model to recognize the construction in one corpus and seeing if it succeeds in another. The stabilizers that maintain this pattern include its frequency of use, the redundancy of its multiple cues (if one fails, others compensate), and normative enforcement through editorial practices that correct malformed instances. Equally important are the failure cases – categories that don’t qualify as HPC kinds. Some proposals are too thin: one-off items (like nonce coinages such as bromance before it caught on), speech errors that blend words together, and children’s overregularizations (goed for ‘went’) lack the stabilizers that would make them predictable beyond their immediate context. Others are too fat: when typologists group together all “resultative” constructions across languages, or all “ditransitive” patterns, they’re pooling structures maintained by largely distinct mechanisms in each language – different morphosyntactic resources, cue reliabilities, and normative regimes (Haspelmath 2010). With many mechanisms local to each language, the cross-linguistic umbrella groups multiple language-internal kinds rather than forming a single HPC kind. Finally, negative or complement classes (e.g., “all ungrammatical strings”) are defined by what they’re not, rather than by shared properties held together by identifiable mechanisms. The discipline here is mechanism-first, in line with the word-kinds program (Miller 2021). The framework yields predictions and disconfirmers. A perturbation prediction: weakening a stabilizer (e.g., lowering frequency or impoverishing input) will reduce cluster covariance before norms re-stabilize – a pattern testable under register shifts or in learner corpora. A scaling prediction: rarer, articulatorily marked segments that lack quantal robustness (e.g., /y/ in the vowel space, which requires precise front-rounding coordination) become more probable as system size grows; analogously, low-frequency constructional variants should be more prevalent in larger constructicons. Both predictions borrow directly from the phoneme results (Ekström et al. 2025: Fig. 2 p. 7) and extend them to higher levels. In sum, properly operationalized HPC naturalizes linguistic ontology for cognitive science. It tells us when a category is the right sort of thing to underwrite inference, what keeps it 6
stable enough, and how it can change. Each tier has its stabilizers; each stabilizer leaves signatures; each signature supports prediction. The result isn’t that everything is an HPC, but that many linguistic categories pass a disciplined test – and some don’t. 1 Framework and diagnostics The introduction outlined how linguistic categories might qualify as homeostatic propertycluster kinds. This section operationalizes that claim into executable diagnostics with specific thresholds and falsifiable predictions. Claims about kindhood depend on identifiable mechanisms within populations: English /r/ is stabilized by English articulatory norms and community transmission, not by universal phonetic laws. How long a kind persists, and how widely it extends, are empirical questions about whether stabilizing mechanisms maintain covariance. Cross-linguistic categories like “resultative” often pool constructions maintained by completely different mechanisms in each language – these are useful classifications but not single kinds. The framework discovers kinds empirically rather than declaring them universally. The two diagnostics – projectibility and homeostasis – require different evidence at each linguistic level: Boyd treats projectibility as following from homeostatic maintenance: if mechanisms hold a cluster together, the cluster will support induction. I revise this relationship, treating them as independent diagnostics for operational clarity and falsifiability. This isn’t mere operationalization – it’s a methodological commitment. The relationship between mechanism and projection is asymmetric in a specific way. Figure 1shows the structure. 1. Projectibility: Can we predict new data from old? If phonemes form genuine kinds, then knowing a language’s family should constrain expectations about its inventory size. If words are kinds despite semantic drift, then patterns learned from earlier 7
Mechanisms Covariance Projection maintain supports Ontological (causal) Test projection Check mechanisms Warrant kindhood informsif both succeed Epistemological (evidential) Projection without mechanisms: spurious overfitting Mechanisms without projection: broken homeostasis Figure 1: The causal arrow runs from mechanisms to projection; the warrant arrow runs the reverse. Both diagnostics have to succeed independently. decades should help identify the same word later. If constructions are kinds, then cue patterns from one corpus should work in another. By contrast, nonce words or one-off blends would fail this test – they lack the regularities that enable prediction. This isn’t philosophical hand-waving – it means concrete metrics like cross-validated prediction accuracy and held-out F1 scores. 2. Homeostasis: What keeps the patterns stable? The term “homeostasis” captures how mechanisms actively maintain property clusters despite perturbations, analogous to biological self-regulation. Projectibility is an epistemic success criterion (out-of-sample prediction); homeostasis is the ontological claim that mechanisms maintain the relevant property cluster. For each category, I name specific mechanisms and look for their signatures. Phonemes are stabilized by quantal regions that create articulatory “sweet spots” (Stevens 1989), dispersion that maintains distinctiveness (Liljencrants & Lindblom 1972), perceptual magnets that pull varied pronunciations toward prototypes (Kuhl et al. 1991), and community norms that transmit systems across generations. These mechanisms predict specific patterns: languages should converge on similar inventory sizes (not scatter randomly), and marked vowels like /y/ should appear mainly in larger systems (not randomly). A category might project without identifiable mechanisms (overfitting to spurious pat8
terns) or exhibit putative mechanisms without projecting (broken homeostasis). Both diagnostics have to succeed independently for kindhood to be warranted. The protocol is straightforward: identify clustered properties, name stabilizers, test prediction, verify signatures. Boyd’s HPC framework treats projectibility as supporting inductive inference: from known properties to novel properties, from observed instances to unobserved cases. The cross-validated prediction tests I employ – held-out accuracy, F1 scores, cross-corpus transfer – operationalize this criterion: if a category supports reliable out-of-sample prediction, it licenses the kind of inductive generalization Boyd has in mind. Where both tests succeed, the category qualifies as an HPC kind over whatever period and extent mechanisms maintain covariance. Where they fail – no predictive power or no credible mechanisms – the category might be useful for description but doesn’t constitute a kind in this technical sense. Section 6 details the failure modes: categories that are too thin, too fat, or merely negative complement classes. Recent work provides models for what success looks like. Phonemes show remarkably consistent inventory sizes across language families and systematic scaling patterns for marked segments – exactly what we’d expect from the mechanisms described above (Ekström et al. 2025). Words maintain enough coherence through meaning change that their identifying features – spelling, pronunciation, distribution – stay bundled at usable levels (Miller 2021). Negative categories don’t qualify: “all ungrammatical strings” or “exceptions to rule R” are defined by what they lack, not by properties held together by mechanisms. The approach remains neutral about grammatical frameworks – it asks what patterns travel and why, not how to represent them theoretically. 2 Methods: tests, scope, and robustness The ontological commitment is robust – these categories are natural kinds maintained by mechanisms – but the epistemology is disciplined: the tests has to be executable and falsifi9
4 Case B – Words: a positive HPC under drift If homeostatic property–cluster kinds are to do real work beyond phonology, they ought to earn their keep where categories are visibly historical. Words are the hard case and the natural next step. The aim here isn’t to freeze a lexeme at a moment in time, but to show that a word can drift semantically while preserving enough covariation among its properties for inductive use. On the projectibility side, the question is whether held-out decades are predictable from earlier usage; on the homeostasis side, the question is whether there are plausible stabilizers that would make such predictability non-accidental. Miller’s mechanism-first treatment of word-kinds sets the bar: kindhood, if it applies, has to be earned a posteriori by sustained covariation among orthography, phonology, meaning, and distribution in a particular population and time slice (Miller 2021). I take a single English lexeme with documented drift – egregious is a convenient instance – and trace its distributional neighbourhood across bins of time. A full multi-lexeme target/control evaluation (not yet undertaken) is to be reported in Appendix A. The operationalization is deliberately spartan. By “distributional neighbourhood” I mean the set of words that typically appear within a fixed window (here, five tokens) of the target word in large corpora. These neighbourhoods are represented as vectors in semantic space, where proximity captures co-occurrence patterns. Representations are decade-binned skip-gram negative sampling (SGNS) embeddings (window = 5, dim = 300) aligned by orthogonal Procrustes on a 1,000-word anchor vocabulary; targets meet a minimum of ≥200 raw tokens per decade. Contexts are aggregated from large, genre-mixed textual sources; tokens are lemmatized and lower-cased; and the representation of a decade is a smoothed average of its local contexts. No sense inventory is imposed; the question isn’t which senses exist, but whether the family of properties that travel with the word remains coherent enough to support inference. Two simple checks supply the answer. First, a cohesion check: nearest-neighbour structure for egregious remains recognizably organized from one decade to the next, even as the 16
centre of gravity shifts – i.e., drift is visible but not chaotic. For egregious specifically, the nearest neighbours shift from {distinguished,excellent,notable} in the 1900s–1920s to {violation,error,abuse} by the 1990s–2000s, tracking the semantic drift from positive to negative valence. But the intermediate decades show gradual transition rather than abrupt reorganization: the 1950s–1960s neighbourhoods include both {notable,prominent} and {mistake, fault}, capturing the word mid-drift. Cohesion metrics satisfy our thresholds (top-50 neighbour overlap ≥0.30; rank-correlation across decades), with values reported in Appendix A. Distributional neighbourhoods serve as proxies for patterns of use. When sense distinctions matter, these distributional patterns can be cross-checked against explicit semantic criteria (e.g., dictionary senses, human sense annotation). Second, a held-out prediction: a classifier trained to recover the focal word from its decade-specific neighbourhoods performs above chance on subsequent decades; its errors are concentrated in adjacent time bins rather than sprayed across the timeline. The classifier exceeds the shuffled-label baseline by ≥0.10 F1, and its mean absolute temporal error is ≤one decade. For egregious, classification F1 = 0.42 (exceeding our 0.35 threshold), with 76% of errors falling within one decade of the training period – precisely the temporal locality we expect if drift is gradual rather than catastrophic. Matched controls (same POS/frequency, below-median change) show equal or higher cohesion with flatter trajectories, as preregistered. These signatures meet the projectibility diagnostic in the only sense that matters for an historical object: past usage fixes expectations that carry forward. The relevant test is whether the distributional neighbourhood resists dissolution when the mean location of a word’s use shifts; that resistance is precisely what the HPC picture predicts when stabilizers preserve enough shared properties as a word moves. Those stabilizers are not mysterious. Orthographic standardization constrains spelling through educational institutions and publishing practices, implemented cognitively via explicit instruction and error correction that builds orthographic representations resistant to variation. Frequency-based entrench17
ment makes high-frequency items resistant to perturbation through sheer repetition in memory: each token strengthens the form-meaning link and automates retrieval (Bybee & Hopper 2001). Editorial norms are enforced through copy-editing workflows that flag nonstandard usage, implemented via conformity bias (speakers align with prestigious variants) and reputation monitoring (fear of correction motivates norm-following). Register licensing operates through genre conventions that sanction certain words in certain contexts, implemented via associative learning linking words to situational contexts. These are not mere labels for correlation – they are social practices with cognitive effects, operating through domain-general mechanisms like associative memory, conformity, and error-driven learning.3 4.1 Objections and responses There are obvious objections, and they can be separated from one another. One is methodological: distributional neighbourhoods are proxies, not senses. That’s correct, but it isn’t a defect here. The claim under test is that the relevant family of properties stays bundled tightly enough for prediction; distributional stability is an appropriate read-out of that bundling, and the failure mode – a collapse in cohesion and in held-out performance – is clear. A second objection distinguishes homeostasis from inertia: maybe distributional stability reflects momentum (frequency begets frequency) rather than active maintenance. This is testable. Homeostasis predicts that weakening a stabilizer (e.g., reducing editorial oversight in informal registers, removing spelling instruction) should degrade coherence before restabilizing at a new equilibrium. Inertia predicts monotonic decay. The dictionary lag 3The dictionary record illustrates the point. Positive-valenced egregious is attested early in English, but ordinary contemporary usage is dominated by a negative sense (‘conspicuously bad’); major dictionaries show the lag between change in use and lexicographic re-weighting – Webster’s 1828 entry still foregrounds the older, positive sense (Webster 1828), while modern Merriam-Webster lists the negative sense first and marks the earlier sense archaic (Merriam-Webster, Inc. 2025); the OED’s later supplements similarly record and then consolidate the pejorative distribution (egregious, adj. 2025). That chronology – early attestation, community shift, later lexicographic re-ordering – is the kind of covariance Miller diagnoses as constitutive of word-kinds: what matters is the ensemble of interacting stabilizers that carry a word’s cluster of properties forward, not an immutable essence (Miller 2021). 18
for egregious – where prescriptive sources resisted the pejorative sense for decades despite widespread usage – demonstrates active normative pressure, not passive momentum. If stability were mere inertia, lexicographic and colloquial distributions would track together; the lag shows independent stabilizers operating at different speeds. A third objection is genealogical: some drifts are punctuated, driven by contact or fashion, and so prediction should break. This is a fair disconfirmatory case; it’s also consistent with the framework. If a lexeme’s covariance collapses or becomes unmoored from any credible stabilizer, we should withhold kindhood for that population–time slice rather than force an HPC verdict. A fourth objection appeals to polysemy: if a family of uses fractionates, does the umbrella still count as one kind? Here again the framework is conservative. Population-relative equilibria are what matter. If distinct usage communities stabilize distinct covariances – two clusters with their own stabilizers – the correct description is two local kinds, not an analyst’s disjunctive lump. The positive case, then, is modest but informative. Words can change and still remain HPC kinds because the mechanisms that matter for use – orthography and phonology, frequency and register, collocational habits and editorial standards – are strong enough to bind their properties across time. What the figures show isn’t that egregious means now what it meant in older prose, but that its empirical profile remains predictively structured as it moves. That’s exactly the sense in which a linguistic category earns kindhood by HPC lights: it projects because it’s homeostatically maintained. The contrast class is equally clear. One-off coinages that fail to diffuse, campaign-season blends, or transient vogue terms often show neither cohesion nor predictive grip when tracked across time; nothing stabilizes them. They are legitimate objects of study, but they aren’t kinds in the relevant sense. This tier, then, completes the bridge from phonemes to higher structure. In phonology, the stabilizers are largely biophysical and perceptual and yield stability bands and scaling relations. In the lexicon, the stabilizers are mainly sociocultural and distributional and yield 19
cohesion under drift and recoverable neighbourhoods. The ontology is the same: a posteriori kinds stabilized by mechanisms we can name, whose signatures we can see. 5 Case C – Constructions: let alone as a positive HPC Constructions – conventionalized pairings of form and meaning that go beyond compositional rules – offer a different challenge for the HPC framework than single segments or words. Where phonemes cluster in articulatory space and words maintain distributional neighbourhoods, constructions rely on multiple converging cues that speakers recognize as a gestalt. The English let alone construction provides an ideal test case: it has clear formal markers, well-studied semantics, and occurs frequently enough in corpora to enable quantitative analysis. Consider the contrast in (1): (1) a. I can’t afford coffee, let alone dinner. b. *I can afford coffee, let alone dinner. This requires a negative or scalar context (1a) and signals that the second item is even less likely than the first. Following Fillmore, Kay & O’Connor (1988), this scalar relationship – where Yranks higher than Xon some contextually relevant scale – defines the construction’s core meaning. But how do speakers recognize this pattern reliably? And what keeps its formal and semantic properties bundled together across different texts and registers? I profile three observable cues that work together: •String anchor: The phrase let alone itself •Syntactic parallelism: The contrasted elements (Xand Y) typically match in grammatical category (both nouns, both verb phrases, etc.) •Licensing context: negative/downward-entailing markers such as not,n’t,no never, 20
hardly,without,even that reverse normal entailment patterns4 What is this construction cognitively? I treat let alone as a stored form-meaning pairing: a schema specifying (a) the anchor string, (b) a syntactic template [X, let alone Y] with parallelism expectations, and (c) scalar semantics (Y more extreme than X in a downwardentailing context). The schema abstracts over exemplar instances. This is compatible with usage-based Construction Grammar (Goldberg 1995,2006), Sign-Based Construction Grammar (SagEtAl2012), HPSG (PollardSag1994), or any framework treating constructions as stored units rather than purely derived structures. The commitment is to storage + form-meaning pairing; the specific theory of how acquisition and processing implement this is an open question. For concreteness, I assume cue weights emerge from frequency and reliability (usage-based mechanisms), but constraint-based or parametric implementations could maintain the cluster via different stabilizers. To test whether this cue bundle qualifies as homeostatic, I use two independently annotated corpora from the Universal Dependencies (UD) project: English GUM (Zeldes 2017) – containing diverse genres from academic writing to online reviews – and English EWT (Silveira et al. 2014) – built from web text and email. These corpora provide syntactic parses that identify grammatical categories and dependencies, enabling automatic extraction of the parallelism and licensing features beyond simple string matching. The projectibility test asks: can patterns learned in one corpus predict instances in another? I train a minimal classifier on the three-cue bundle using data from GUM, then evaluate its ability to identify true let alone constructions in the held-out EWT corpus (and vice versa). This cross-corpus design is crucial – if the construction were merely a frozen idiom or a corpus-specific quirk, the patterns wouldn’t transfer. Table 1reports the discrimination performance using PR-AUC. Because evaluation is restricted to anchor-present candidates, the anchor-only baseline is by construction uninformative (PR–AUC ≈0.50), so gains reflect genuine cue structure. 4In downward-entailing contexts, inferences flip: from can’t afford dinner we can infer ‘can’t afford ex21
Table 1: Cross-corpus evaluation for let alone. Full model uses anchor+parallelism+licensing; ablations drop one cue. Matched decoys, balanced classes. Sample sizes: GUM n= 12, EWT n= 15 (counts reflect the anchor-present evaluation set). Direction Model PR-AUC ∆ GUM→EWT Full bundle 0.750 – Drop parallelism 0.600 –0.150 Drop licensing 0.575 –0.175 EWT→GUM Full bundle 0.725 – Drop parallelism 0.500 –0.225 Drop licensing 0.525 –0.200 Note: All models achieve Recall = 0.500 by design (balanced evaluation). F1 scores range from 0.500–0.600. Full precision/recall/F1 values available in supplementary materials. Given the small anchor-present counts (GUM n= 12, EWT n= 15), these results are illustrative; the confirmatory analysis will expand to additional UD English corpora and a larger anchor sweep. The full three-cue model achieves PR-AUC ≥0.70 in both transfer directions (GUM→EWT means training on GUM, testing on EWT). This substantially exceeds a shuffled-label baseline and indicates robust cross-corpus generalization. More tellingly, removing either the parallelism cue or the licensing cue degrades performance by ∆≥0.10 – exactly what we’d expect if these features work together homeostatically. The string anchor alone isn’t sufficient; the construction needs its supporting cast of cues. Figure 4decomposes the cue distributions in each corpus. Despite different genres and collection methods, both corpora show remarkably similar profiles: parallelism rates hover around 80%, verbs and nouns dominate the Yposition, and licensing elements appear in roughly 60% of cases. This cross-corpus stability – maintained without explicit coordination between the corpus creators – suggests genuine linguistic regularities rather than annotation artifacts. What mechanisms maintain this stability? Three are plausible and testable: 1. Frequency and entrenchment: The construction appears often enough (dozens of times even in modest corpora) that speakers internalize its pattern through repeated exposure pensive dinner’, whereas from can afford dinner we cannot infer ‘can afford expensive dinner’. 22
GUM EWT 0.0 0.2 0.4 0.6 0.8 1.0 Parallelism rate N=12 N=15 Parallelism GUM EWT 0.0 0.2 0.4 0.6 0.8 1.0 Proportion of Y heads Distribution of Y head UPOS UPOS VERB NOUN ADJ OTHER GUM EWT 0.0 0.2 0.4 0.6 0.8 1.0 Licensing prevalence Licensing Figure 4: Cue profile for the let alone construction in UD English GUM and EWT. Left to right: proportion of tokens with syntactic parallelism (matching Universal Part-of-Speech (UPOS) match between contrasted heads for the heads of Xand Y), distribution of Y–head UPOS, and prevalence of scalar/downward–entailing licensing items within five tokens to the left. Error bars are bootstrap 95% intervals (2,000 resamples). Token counts and exact estimates are reported in the repository tables. 2. Cue redundancy: Multiple signals converge – even if parallelism fails in a rushed email, the anchor and licensing context still signal the construction 3. Normative pressure: Editorial practices and style guides reinforce the canonical pattern, especially in formal registers; malformed instances like *I bought coffee let alone dinner would likely be corrected in editing These mechanisms operate at different timescales but interact to co-stabilize the construction. Frequency and entrenchment work rapidly (milliseconds to weeks): each token use strengthens memory traces, increasing production probability, which generates more tokens in a self-reinforcing loop (Bybee & Hopper 2001). Editorial norms and genre licensing operate slowly (years to decades): copy-editing workflows and style guide updates formalize emergent patterns, creating explicit standards that then constrain future production. The fast loop generates the pattern; the slow loop crystallizes and transmits it. Perturbation experiments can distinguish their contributions: reducing frequency while maintaining editorial standards should weaken the cluster gradually, whereas removing editorial oversight while maintaining frequency should increase drift without collapse. Both stabilizers are necessary: 23
0.0 0.2 0.4 0.6 0.8 1.0 Recall 0.0 0.2 0.4 0.6 0.8 1.0 Precision PR curves: train GUM test EWT Full bundle No parallelism No licensing Figure 5: Projectibility and ablation for let alone. Precision–recall curves for a regularized logistic model using the full cue bundle (anchor + parallelism + licensing; solid) versus ablations (drop parallelism; dashed; drop licensing; dotted). Model is trained on GUM and evaluated on EWT; class prevalence in the target set and train/test construction are held constant across conditions. Shaded bands are bootstrap 95% intervals. The full bundle achieves high PR–AUC (≥0.70) and each ablation reduces PR–AUC by at least 0.10, consistent with a homeostatically maintained cue bundle. frequency alone produces transient patterns (internet slang that fades quickly), normative pressure alone produces rigid prescriptions that speakers ignore (failed language reforms). The interaction explains the observed robustness. These mechanisms – frequency, redundancy, and normativity – are precisely the homeostatic forces that Boyd’s framework predicts for socially maintained kinds. Unlike the biophysical constraints that stabilize phonemes, constructions rely more heavily on usage-based learning and community standards. But the empirical signatures are parallel: predictable patterns that degrade systematically when stabilizers are removed. 24
How do speakers acquire this cue bundle? Children learning let alone face a distributional learning problem: from scattered instances in input, extract the recurring pattern that the anchor co-occurs with parallelism, licensing context, and scalar meaning. Domain-general statistical learning mechanisms – tracking form-meaning co-occurrences, registering cue reliability, chunking frequent sequences – are sufficient (Tomasello 2003,Goldberg 2006). As the construction becomes entrenched through repeated activation, it gains processing advantages (faster recognition, automatic retrieval) that further stabilize it against perturbation. This is the cognitive implementation of “frequency and entrenchment”: repeated exposure →strengthened memory trace →resistance to change. Frequency drives entrenchment through repeated activation of form-meaning mappings in memory (Bybee & Hopper 2001); cue redundancy provides fallback signals when individual cues are noisy or degraded; normative pressure operates through explicit correction (editorial changes, teaching) and implicit modelling (exposure to edited text). The ablation results in Table 1provide evidence that these mechanisms matter: removing parallelism or licensing degrades performance because the construction depends on their combined contribution. The small sample sizes (GUM n= 12, EWT n= 15) mean these results are illustrative rather than definitive; general claims about constructions-as-kinds will require broader sampling across construction types and corpora. What let alone demonstrates is proofof-concept: the diagnostics can be applied to constructions, and they yield interpretable signatures when they succeed. A further question concerns constructional inheritance. Let alone isn’t isolated – it shares properties with a family of scalar additive constructions: much less,not to mention,never mind,to say nothing of. All require downward-entailing licensing and signal scalar extremity, but differ in register (formal vs. colloquial), cue reliability (e.g., much less shows weaker parallelism), and productivity. An HPC treatment of this family would ask: is the scalaradditive schema itself an HPC kind at a higher level of abstraction, with let alone as an instantiation? Or are these distinct kinds that happen to share features? The diagnostics 25
have little to say about drift, diversity, and social maintenance. The present approach keeps the realism while naturalizing it: kinds are whatever supports reliable inference because stabilizers – biophysical, developmental, social – keep enough of the relevant properties together (Miller 2021,Boyd 1991,1999). The figures and tables in this paper are small demonstrations of that general point. Where the signed effects and thresholds are met, a kind claim is warranted; where they aren’t, the label should be withheld. That discipline, and not any particular representation, is the contribution. 32
A Statistical specifications and robustness checks This appendix provides complete technical specifications for the analyses in the main text. All thresholds and decision rules were fixed before analysis. A.1 Phoneme-level specifications Model specification. The /y/ presence model uses logistic regression with fixed effects for language family and macro-area (dummy codes) plus centred vowel inventory size. Cross-validation. 10-fold CV using GroupKFold by language family to prevent leakage across related languages. Evaluation metric: ROC-AUC (area under receiver operating characteristic curve). Success criteria. All have to be met: (1) positive inventory effect with 95% CI excluding zero; (2) 10-fold CV AUC ≥0.70; (3) Mann–Kendall trend test p < 0.01. Trend significance is computed with a Mann–Kendall-style statistic (normal approximation) and a permutation null (1000 permutations); (4) effect persists across three specifications: family-effects-only, family+area effects, and lineage-pruned samples (all including the inventory predictor). A.2 Word-level specifications Target selection. Top decile of diachronic change scores from Hamilton, Leskovec & Jurafsky (2016), filtered to maintain ≥200 raw tokens per decade (typically 1–100 per million). Controls matched on POS and log-frequency (±0.5) within the same decade windows, but below median change score. A.3 Construction-level specifications Features. (1) Anchor: binary presence; (2) Parallelism: UPOS match between contrasted heads; (3) Licensing: presence of not,-n’t,no,never,hardly,without,even within 5 tokens left. 33
Model and evaluation. L2-regularized logistic regression (C=1.0). Evaluation restricted to anchor-present candidates (true let alone vs. strings containing “let alone” but failing syntactic/semantic criteria). Sample sizes in Table 1reflect the anchor-present evaluation set. Success criteria. (1) Cross-corpus PR-AUC ≥0.70; (2) Each ablation (drop parallelism; drop licensing) reduces PR-AUC by ≥0.10; (3) Anchor-only baseline on the anchor-present evaluation set behaves as expected (≈0.50 PR-AUC); (4) Calibration slope 0.8–1.2, intercept ±0.2; (5) Performance exceeds shuffled-label baseline. A.4 Multiple testing and inference Primary outcomes (no correction). /y/ slope and AUC (Case A); average cohesion and F1 across target/control pairs (Case B); cross-corpus PR-AUC and mean ablation delta (Case C). Secondary outcomes. Benjamini–Hochberg correction at q= 0.10 for: (1) individual family medians; (2) multiple vowel comparisons; (3) individual lexeme metrics. Uncertainty. Bootstrap CIs (2000 resamples) for: family medians, AUC metrics, classification metrics. Permutation tests (1000 permutations) specifically for Mann–Kendall trend statistics. A.5 Perturbation experiments Frequency downsampling. Reduce construction tokens by 75%, 50%, 25% via stratified sampling. Recompute cue covariance (phi coefficients) and PR-AUC. Success: ≥0.10 drop in PR-AUC or ≥20% reduction in parallelism/licensing rates at 25% sample. Constructicon scaling. Bin corpora by construction type count (quartiles). Estimate P(rare variant) per bin with Wilson CIs. Success: monotonic increase with non-overlapping CIs for extreme quartiles. 34
A.6 Software and versions R 4.3.1 (phoneme analyses): lme4 1.1-34, ggplot2 3.4.2, boot 1.3-28. Python 3.10.12 (word/construction): scikit-learn 1.3.0, pandas 2.0.3, numpy 1.24.3, gensim 4.3.1, stanza 1.5.0. All random seeds fixed at 42. Complete session info in repository SESSION.txt. 35
References Boyd, Richard N. 1988. How to be a moral realist. In Geoffrey Sayre-McCord (ed.), Essays on moral realism, 181–228. Ithaca, NY: Cornell University Press. Boyd, Richard N. 1991. Realism, anti-foundationalism and the enthusiasm for natural kinds. Philosophical Studies 61(1). 127–148. https://doi.org/10.1007/BF00385837. Boyd, Richard N. 1999. Homeostasis, species, and higher taxa. In Robert A. Wilson (ed.), Species: new interdisciplinary essays, 141–186. Cambridge, MA: MIT Press. Boyd, Richard N. 2000. Kinds as the “workmanship of men”: Realism, constructivism, and natural kinds. In Julian Nida-Rümelin (ed.), Rationalität, realismus, revision / rationality, realism, revision: proceedings of the 3rd international congress of the gesellschaft für analytische philosophie, 52–89. Berlin: De Gruyter. https://doi.org/10.1515/ 9783110805703.52. Bybee, Joan L. & Paul J. Hopper. 2001. Introduction to frequency and the emergence of linguistic structure. In Joan L. Bybee & Paul J. Hopper (eds.), Frequency and the emergence of linguistic structure (Typological Studies in Language 45), 1–24. Amsterdam: John Benjamins. https://doi.org/10.1075/tsl.45.01byb. Croft, William. 2001. Radical construction grammar: syntactic theory in typological perspective. Oxford: Oxford University Press. egregious, adj. 2025. Oxford English Dictionary.https://www.oed.com/entry/egregious (13 October, 2025). Ekström, Axel G., Claudio Tennie, Steven Moran & Caleb Everett. 2025. The phoneme as a cognitive tool. Topics in Cognitive Science. Advance online publication. https : //doi.org/10.1111/tops.70021. Fillmore, Charles J., Paul Kay & Mary Catherine O’Connor. 1988. Regularity and idiomaticity in grammatical constructions: The case of let alone.Language 64(3). 501–538. 36
Goldberg, Adele E. 1995. Constructions: A construction grammar approach to argument structure (Cognitive Theory of Language and Culture). Chicago: University of Chicago Press. Goldberg, Adele E. 2006. Constructions at work: The nature of generalization in language. Oxford: Oxford University Press. https://doi.org/10.1093/acprof:oso/9780199268511. 001.0001. Hamilton, William L., Jure Leskovec & Dan Jurafsky. 2016. Diachronic word embeddings reveal statistical laws of semantic change. In Proceedings of the 54th annual meeting of the association for computational linguistics, 1489–1501. Berlin. https://doi.org/10. 18653/v1/P16-1141. Haspelmath, Martin. 2010. Comparative concepts and descriptive categories in crosslinguistic studies. Language 86(3). 663–687. https://doi.org/10.1353/lan.2010.0021. Heyes, Cecilia. 2018. Précis of Cognitive Gadgets: The Cultural Evolution of Thinking. Behavioral and Brain Sciences 42. https://doi.org/10.1017/S0140525X18002145. Khalidi, Muhammad Ali. 2013. Natural categories and human kinds: Classification in the natural and social sciences. Cambridge: Cambridge University Press. https://doi.org/ 10.1017/CBO9781139207071. Kuhl, Patricia K., Karen A. Williams, Francisco Lacerda, Kenneth N. Stevens & Björn Lindblom. 1991. Human adults and human infants show a “perceptual magnet effect” for the prototypes of speech categories, monkeys do not. Perception & Psychophysics 50(2). 93–107. https://doi.org/10.3758/BF03212211. Liljencrants, Johan & Björn Lindblom. 1972. Numerical simulation of vowel quality systems: The role of perceptual contrast. Language 48(4). 839–862. https://doi.org/10.2307/ 411991. Lindblom, Björn. 1990. Explaining phonetic variation: A sketch of the H&H theory. In William J. Hardcastle & Alain Marchal (eds.), Speech production and speech modelling, 37
403–439. Dordrecht: Kluwer Academic Publishers. https://doi.org/10.1007/978-94009-2037-8_16. Manning, Benjamin S. & John J. Horton. 2025. General social agents. Most recent draft: 2025-09-03. Manuscript, MIT; MIT & NBER. Merriam-Webster, Inc. 2025. Egregious.https://www.merriam-webster.com/dictionary/ egregious. Accessed 13 Oct 2025. Miller, J. T. M. 2021. Words, species, and kinds. Metaphysics 4(1). 18–31. https://doi. org/10.5334/met.70. Moran, Steven, Daniel McCloy & Richard Wright (eds.). 2019. PHOIBLE 2.0.https:// phoible.org. Accessed 2025-09-01. Jena. Rubin, Michael. 2008. Is goodness a homeostatic property cluster? Ethics 118(3). 496–528. Silveira, Natalia, Timothy Dozat, Marie-Catherine de Marneffe, Samuel R. Bowman, Miriam Connor, John Bauer & Christopher D. Manning. 2014. A gold standard dependency corpus for English. In Proceedings of the ninth international conference on language resources and evaluation (lrec’14). Reykjavik. Stevens, Kenneth N. 1989. On the quantal nature of speech. Journal of Phonetics 17. 3–45. Tomasello, Michael. 2003. Constructing a language: A usage-based theory of language acquisition. Cambridge, MA: Harvard University Press. Webster, Noah. 1828. An American dictionary of the English language. Entry: “egregious.” Digitized editions available online. Springfield, MA: S. Converse. Zeldes, Amir. 2017. The GUM corpus: Creating multilayer resources in the classroom. In Proceedings of the 2017 conference on empirical methods in natural language processing: system demonstrations. Copenhagen. 38