Full text
! Dissertation an der Universität Osnabrück zur Erlangung des Doktorgrades Dr. rer. nat. Thema The role of language and pragmatics in conceptual abstraction: Interactive experiments and emergent communication models using reference games Kristina Kobrock Die!Dissertation!wurde!betreut!durch Prof.!Dr.!Nicole!Gotzner! Prof.!Dr.!Elia!Bruni! Universität Osnabrück Fachbereich Humanwissenschaften Kognitionswissenschaft Osnabrück 2025
blue Agents are embedded in an environment constructed by the researcher. It is this environment that determines the social interactions of agents and also what the agents will communicate about. — Christiansen & Kirby (2003, p. 304) This document was typeset with the help of KOMAScript and LaTeX using the kaobook class.
The role of language and pragmatics in conceptual abstraction Interactive experiments and emergent communication models using reference games Kristina Kobrock
Acknowledgements Interaction is not only one of the core topics in this dissertation, but it also played a significant role in its development process. During the time I was working on this dissertation, I had the pleasure of interacting with many great scientists, amazing personalities, and wonderful colleagues. First and foremost, I would like to wholeheartedly thank my supervisors, Nicole Gotzner and Elia Bruni, for embarking on this journey with me and for their continuous support over the past three years. I am grateful to you for being open to and encouraging the research project, for helping me grow personally and academically, and for always asking me the right challenging questions and giving me critical feedback when I needed it. Your support has been invaluable to me on my way to becoming an independent researcher! Nicole, you have integrated me into the SPA lab and shown me how supportive and welcoming academia can be. Not only are you an exceptional academic supervisor, but I am also very grateful to you for believing in me, encouraging me, and celebrating with me. Thank you for giving me security amid the precarity of academic positions and offering me the opportunity to continue our joyful and fruitful collaboration! Many thanks also to all members of the SPA lab who have accompanied me, current and former: Charlotte Uhlemann, Elli Tourtouri, Morwenna Hoeks, Radim Lacina, Katja Ruge, Julia Reuter, and Stavroula Alexandropoulou. Especially to Charlotte, who accompanied me through the challenges of developing and implementing a novel experimental paradigm, who was the first student I ever supervised, and who has become my friend, office mate, co-author, and road trip buddy. Special thanks also belong to Xenia Ohmer, who was like an academic big sister to me. Thank you for taking the time for our regular meetings and discussions. I learned a lot from you, and I could always count on your thoughtful advice, timely feedback, keen eye for small mistakes, and creative ideas! I would also like to express my gratitude to the people at the Institute of Cognitive Science who contribute to this being such an ideal place for this academic journey. In particular, I would like to thank all the members of the Computational Cognition research training group for the stimulating scientific discussions, the helpful career advice, and for making me feel that I am not alone with my doubts and problems. Another great source of inspiration and joy during this journey has been working with the amazing Coxi students. Thank you, Ilva Hovemann, Anna Briotto, Eosandra Grund, Vanessa Verdugo, Muhip Tezcan, Vanessa Obi, Felix Japtok, Isabella del Pozo, and Mariia Gudkova, for supporting this research as a student assistant or while writing your thesis with me. I could always learn something new while working with you, and I wish you all the best for your future! My own journey as a Coxi was made special by my dear friends Julian, Johanna, Jana, Merri, Daniel, and Klara (not all of whom are Coxis, but they all were for at least one day).
There are many more people with whom I have enjoyed delightful times at conferences, relaxing and stimulating lunch and coffee breaks, and fun DnD evenings at the institute. Thank you all very much! Whenever I needed a break, a hug, or a listening ear, my family and friends were there for me. I want to thank my parents and my sister. For as long as I can remember, you have looked after me, advised and motivated me, supported me, and loved me. Thank you to my partner Vincent, who distracted me from work when I couldn’t stop, helped me relax when I was stressed, and for your unconditional love and support. Anne, thank you for being my best friend and my constant conversation partner since first grade. Last but not least, I owe a huge thank you to my pre-PhD collaborators and mentors: Timo Röttger, Michael Franke, Iris van Rooij, Gordon Pipa, and Peter König. Without you, none of this would have been possible.
Abstract Choosing the right word to convey a particular concept is one of the central decisions in communication, as natural languages offer multiple ways to communicate the same object (e.g., dalmatian, dog,animal). Such terms reflect different conceptual perspectives, and they can be situated within a conceptual hierarchy, ranging from subordinate (dalmatian) to superordinate (animal) terms (E. V. Clark, 1997; Rosch et al., 1976). The primary objective of this dissertation is to investigate conceptual abstraction, that is, the abstraction process involved in forming and communicating concepts that lie on different levels of a conceptual hierarchy. To this end, I present four case studies that collectively provide insights into answering the following overarching research question: “What is the role of language and pragmatics in conceptual abstraction?” We employ an interactive, reference-game-based experimental paradigm in case study 1 and a computational modeling approach in case studies 2-4 to address this question. The computational approach models language evolution and emergence through repeated interactions between Artificial Neural Network agents, where we manipulate the communicative need via conceptual and contextual factors. In case study 1, we find that cognitive and pragmatic factors play a role in the communication of concepts and that speakers, but not listeners, benefit from the production of cognitively economical terms. Case study 2 demonstrates that the availability of context during communication shapes an emerging language to the extent that it becomes more efficient and less overinformative. Case study 3 extends on these insights and generates three main research findings. First, context-based pragmatics leads to the emergence of very efficient languages. Second, utility-based pragmatics improves efficiency further, but only if languages evolved under the condition of context-based pragmatics. Third, the structure of categories and linguistic systems can both be explained by efficiency, operationalized as a tradeoff between the speaker’s need for simplicity and the listener’s need for informativeness during communicative interaction. Finally, case study 4 probes which linguistic strategies artificial agents spontaneously use when they need to communicate concepts at a higher or lower level of the conceptual hierarchy than those conventionalized in their languages. This study shows that generalization to lower levels of the conceptual hierarchy is based on a compositional strategy, whereas abstraction to higher levels of the conceptual hierarchy benefits from a meaning extension approach. Based on these findings, I argue that language reflects the structure of conceptual hierarchies due to their shared origin in interaction. Pragmatics shapes these linguistic interactions during language use and during language emergence, thereby influencing conceptual abstraction. Moreover, novel communicative needs can be met by using creative linguistic strategies and utility-based pragmatic reasoning, preserving the flexibility of language. The here presented work has implications for current and future research in Cognitive Science, Linguistics, and Pragmatics by demonstrating that interaction and efficiency critically determine conceptual abstraction, by providing first evidence for a joint evolution of efficient language and category systems through communicative interaction between speakers and listeners, and by offering ways forward for the research on the role of pragmatics in conceptualization and lexical choice.
Zusammenfassung Die Wahl des passenden Wortes zur Übermittlung eines bestimmten Konzepts stellt einen der zentralen Entscheidungsprozesse in der Kommunikation dar, denn natürliche Sprachen bieten mehrere Möglichkeiten, dasselbe Objekt zu kommunizieren (z. B. Dalmatiner,Hund,Tier). Die Begriffsvarianten spiegeln unterschiedliche konzeptuelle Perspektiven wider und können innerhalb einer Begriffshierarchie angesiedelt werden, die von untergeordneten Begriffen (Dalmatiner) zu übergeordneten Begriffen (Tier) reicht (E. V. Clark, 1997; Rosch et al., 1976). Das Hauptziel dieser Dissertation ist es, die konzeptuelle Abstraktion, d. h. den Abstraktionsprozess, der bei der Bildung und Kommunikation von Konzepten auf verschiedenen Ebenen einer Begriffshierarchie stattfindet, zu untersuchen. Zu diesem Zweck geben vier Fallstudien gemeinsam Einblicke in die Beantwortung der folgenden übergeordneten Forschungsfrage: „Welche Rolle spielen Sprache und Pragmatik in der konzeptuellen Abstraktion?“ Zur Beantwortung dieser Frage finden ein interaktives, auf Referenzspielen basierendes Experiment in Fallstudie 1 und ein computergestützter Modellierungsansatz in den Fallstudien 2-4 Anwendung. Der computergestützte Ansatz modelliert Sprachentwicklung und -entstehung durch wiederholte Interaktionen zwischen als künstliche neuronale Netze implementierten Agenten, wobei die Kommunikationsbedürfnisse über konzeptionelle und kontextuelle Faktoren manipuliert werden. In Fallstudie 1 stellt sich die Erkenntnis heraus, dass kognitive und pragmatische Faktoren bei der Kommunikation von Konzepten eine Rolle spielen und dass Sprecher*innen davon profitieren, kognitiv ökonomische Begriffe zu wählen, jedoch nicht ihre Zuhörer*innen. Fallstudie 2 zeigt, dass die Verfügbarkeit von Kontext während der Kommunikation eine entstehende Sprache insoweit prägt, dass sie effizienter und weniger überinformativ gestaltet wird. Fallstudie 3 baut auf diesen Erkenntnissen auf und liefert drei wesentliche Forschungsresultate. Erstens führt kontextbasierte Pragmatik zur Herausbildung sehr effizienter Sprachen. Zweitens verbessert nützlichkeitsbasierte Pragmatik die Effizienz weiter, jedoch nur, wenn sich Sprachen unter den Bedingungen kontextbasierter Pragmatik entwickelt haben. Drittens lassen sich sowohl die Struktur von Kategorien als auch die Struktur von Sprachsystemen durch Effizienz erklären, die als Kompromiss zwischen dem Bedürfnis der Sprecher*innen nach Einfachheit und dem Bedürfnis der Zuhörer*innen nach Informativität während der kommunikativen Interaktion operationalisiert werden kann. Schließlich ermittelt Fallstudie 4, welche sprachlichen Strategien künstliche Agenten spontan anwenden, wenn sie Konzepte auf einer höheren oder niedrigeren Ebene der Begriffshierarchie kommunizieren müssen als diejenigen, die in ihren Sprachen konventionalisiert sind. Diese Studie lässt die Schlussfolgerung zu, dass die Verallgemeinerung auf niedrigere Ebenen der Begriffshierarchie auf einer kompositionalen Strategie basiert, während die Abstraktion auf höhere Ebenen der Begriffshierarchie von einem Ansatz der Bedeutungserweiterung profitiert. Diese Erkenntnisse führen zu der Argumentation, dass sich die Struktur von Begriffshierarchien in der Sprache findet, da diese ihren gemeinsamen Ursprung in der Interaktion haben und dem Effizienzdruck unterliegen, der sich aus dem Bedürfnis der Sprecher*innen nach einfacher Sprache und dem Bedürfnis der Zuhörer*innen nach informativen Botschaften ergibt. Die Pragmatik prägt
6.9 Context-aware: Mean entropy scores across all datasets for different context conditions indicated by the number of shared attributes ........................ 63 6.10 Context-unaware: Most errors occur on the diagonal from top left to bottom right, i.e. in the finest possible context conditions ............................. 67 6.11 Context-aware: Most errors occur on the diagonal from top left to bottom right, i.e. in the finest possible context conditions ............................... 68 6.12 Context-unaware: Most false negative errors occur in the top left, i.e. in conditions where generic concepts need to be discriminated in fine contexts ................. 69 6.13 Context-aware: Most false negative errors occur in the top left, i.e. in conditions where generic concepts need to be discriminated in fine contexts ................. 69 6.14 Context-unaware: Most false positive errors occur in the bottom right, i.e. in conditions where specific concepts have to be discriminated in a fine context . . . . . . . . . . . . 70 6.15 Context-aware: Most false positive errors occur in the bottom right, i.e. in conditions where specific concepts have to be discriminated in a fine context . . . . . . . . . . . . 71 6.16 Context-unaware: Mean NMI scores across all datasets for different concept (# fixed attributes) and context conditions (# shared attributes) .................. 73 6.17 Context-aware: Mean NMI scores across all datasets for different concept (# fixed attributes) and context conditions (# shared attributes) .................. 73 7.1 Model architecture ....................................... 83 7.2 Game scenarios in Experiment 1 ............................... 86 7.3 Game scenarios added in Experiment 2 ........................... 87 7.4 Message lengths per concept hierarchy level ........................ 90 7.5 NMI, consistency and effectiveness scores for each level of the conceptual hierarchy . . 91 7.6 Distribution of message lengths for different levels of the conceptual hierarchy . . . . . 95 7.7 Tradeoff between lexicon size and informativeness ..................... 96 7.8 Zipf’s law like distribution of message frequency plotted for dataset D(4,4) ....... 97 7.9 Zipf’s law like distribution of message length plotted for dataset D(4,4) ......... 97 7.10 Dataset examples ........................................ 109 7.11 Information-theoretic scores and ambiguity in language .................. 112 7.12 Training trajectory plots with training and validation accuracies for one example dataset D(4,4) and both conditions. .................................. 117 7.13 Zipf’s law like distribution of message frequency plotted for each dataset. ....... 119 7.14 Zipf’s law like distribution of message length plotted for each dataset . . . . . . . . . . 120 8.1 Examples for speaker inputs for training and testing in the two zero-shot test conditions “to specific” and “to generic” ................................. 128 8.2 Architecture ........................................... 129 8.3 To specific: Topographic similarity scores calculated on messages from the train and test splits ............................................... 134
8.4 To generic: Topographic similarity scores calculated on messages from the train and test splits ............................................... 135
List of tables 6.1 Datasets with 𝑛attributes and 𝑘values, labeled as 𝐷(𝑛, 𝑘).. . . . . . . . . . . . . . . . 57 6.2 Context-unaware: Unique messages used to refer to a randomly picked specific concept in the D(4,4) dataset over different context conditions ................... 58 6.3 Context-aware: Unique messages used to refer to a randomly picked specific concept in the D(4,4) dataset over different context conditions ..................... 59 6.4 Accuracy means for agents trained in the context-unaware and context-aware setting averaged over five runs with standard deviations. ..................... 66 6.5 Context-unaware: Unique messages used to refer to a randomly picked specific concept in the D(3,4) dataset over different context conditions ................... 71 6.6 Context-aware: Unique messages used to refer to a randomly picked specific concept in the D(3,4) dataset over different context conditions ..................... 71 6.7 Context-unaware: Unique messages used to refer to a randomly picked specific concept in the D(5,4) dataset over different context conditions ................... 72 6.8 Context-aware: Unique messages used to refer to a randomly picked specific concept in the D(5,4) dataset over different context conditions ..................... 72 7.1 Summary of predictions and measures ........................... 88 7.2 Mean accuracies on the training, validation and test datasets ............... 89 7.3 Mean entropy-based scores, i.e. NMI, effectiveness and consistency ........... 90 7.4 Accuracies on the test dataset ................................. 93 7.5 Lexicon sizes .......................................... 95 7.6 Qualitative examples ...................................... 99 7.7 Statistical models fitted with brm for Experiment 1 ..................... 114 7.8 Statistical models fitted with brm for Experiment 2 ..................... 115 7.9 Posterior summaries of the Bayesian hierarchical models predicting the entropy-based scores, i.e. NMI, effectiveness and consistency ....................... 116 7.10 Mean number of training epochs ............................... 117 8.1 Zero-shot test accuracies for both conditions. ........................ 131 8.2 NMI scores for both conditions. ................................ 131 8.3 The ratio of messages and concepts for the test data split 𝑀𝑡𝑒𝑠𝑡/𝐶𝑡𝑒𝑠𝑡 . Percentage of reused and novel messages from the total set of unique messages 𝑀𝑡𝑒𝑠𝑡.......... 134 8.4 To specific: Training and validation accuracies ....................... 139 8.5 To generic: Training and validation accuracies ....................... 139 8.6 To specific: Number of concepts and messages for the test split and their ratio . . . . . 140 8.7 To generic: Number of concepts and messages for the test split and their ratio . . . . . 140
8.8 To specific: One random example for a specific concept from the test data per dataset, the context condition in which it was presented (in number of shared attributes) and the message that was used to refer to the concept ........................ 141 8.9 To generic: One random example for a generic concept from the test data per dataset, the context condition in which it was presented (in number of shared attributes) and the message that was used to refer to the concept ........................ 141 8.10 To specific: Number of unique concepts in each dataset for each dataset split . . . . . . 142 8.11 To generic: Number of unique concepts in each dataset for each dataset split . . . . . . 143
List of abbreviations and terms AI Artificial Intelligence 23,25,28 ANN Artificial Neural Network 22,23,28 Backpropagation 22–25 BOSS Bank of Standardized Stimuli 7,9,34 DL Deep Learning 22–25 EGG Emergence of lanGuage in Games 27 EOS end-of-sequence 26 FF-NN Feed-forward Neural Network 22,25,27 Gradient 23 GRU Gated Recurrent Unit 23,26,55,168 Gumbel-Softmax 23,24 IL Iterated Learning 20,21 LLM Large Language Model 23,27,28,152,159,160,168 LSTM Long Short-Term Memory 23,55 MARL Multi-agent Reinforcement Learning 22–24 ML Machine Learning 22,23,32,152 RE referring expression xi,35,38–41,44–46 Reinforce 23 RL Reinforcement Learning 22–24,27 RNN Recurrent Neural Network 23,25 RSA Rational Speech Acts 5,18–20,30–32,78,121,122,167 SL Supervised Learning 23 Transformer 23,55,168
1 Motivation and outline One fascinating topic linking Cognitive Science and Linguistics lies in the interplay between language and cognition. There are two research directions: one investigates how language influences cognition, and the other investigates how cognition influences language. On the one hand, we can ask: “How does language influence cognitive processes such as perception and categorization, but also higher-level cognitive processes such as reasoning or problem solving?” This question implies that language is not merely a tool for communication but that it also has a role in cognition. On the other hand, we can ask: “How is language shaped by the cognitive system and the environment?” This question targets language evolution and the pressures and constraints that influence an emerging language. 1.1 Language and cognition How language influences perception and categorization is a key topic in the area of research that investigates the interplay between language and cognition. This kind of research is linked to the Sapir-Whorf hypothesis and started out as a view of linguistic determinism (Sapir, 1912; Whorf, 1956). Whorf (1956), facing the richness of linguistic diversity, proposed that language determined perception, categorization, and, ultimately, actions. This strong view of the Whorfian hypothesis was abandoned in the field (Boroditsky, 2001). However, the weaker view that there is some influence of language on perception is, again, a vivid topic of research under the branding ‘linguistic relativity.’ Recent studies investigate how language impacts perception and categorization, for example, in the domains of color (Regier & Kay, 2009; Winawer et al., 2007), space (Majid et al., 2004), time (Boroditsky, 2001; Boroditsky et al., 2011), or kinship categories (Kemp & Regier, 2012). For instance, Winawer et al. (2007) tested English and Russian speakers in a speeded color discrimination task in which participants had to decide which of two color chips matched a target color. Crucially, English and Russian differ in terms of how these languages divide the color spectrum. While English speakers do not discriminate between lighter blues and darker blues, the Russian language makes this distinction compulsory because there exists a color term for lighter blues (“goluboy”) and one for darker blues (“siniy”). The experimental findings revealed a category advantage for speakers of Russian but not English. Russian speakers were faster at categorizing colors when the two blue color chips fell into different linguistic categories than when they were from the same linguistic category. This was taken as evidence that linguistic representations play a role in tasks that require presumably objective perceptual decisions (Winawer et al., 2007). The effect was also replicated with Greek and German speakers and shows robust electrophysiological signatures in the brain (Maier & Abdel Rahman, 2018). While it is far from being uncontroversial, the claim that language influences perception has gained empirical and theoretical support in recent years (Forder & Lupyan, 2019; Lupyan, 2012; Lupyan et al., 2020).
2 1 Motivation and outline Moving from low-level to high-level cognition, this dissertation aims to proceed from the role of language in perception and categorization towards the role of language in abstraction and generalization. Abstraction is one of the prime higher-level cognitive abilities. It is involved in reasoning, problem solving, complex thoughts, and it has close ties to language. While abstraction lies at the core of Cognitive Science (Tenenbaum et al., 2011), surprisingly little empirical research has been dedicated to abstraction until recently. One reason for this might be the lack of an integrative definition across fields (Burgoon et al., 2013). In the following, I draw on definitions and characterizations provided by Yee (2019) and Burgoon et al. (2013) to elucidate how abstraction can be defined, and what makes abstraction special for cognition. While the notion of abstraction has been used in different manners in the literature, one aspect appears to be central to all views (Burgoon et al., 2013). Abstraction is the process by which we aggregate individual experiences with objects, actions, events, and ideas. 1 We abstract over individual (sensory or motor) experiences and highlight what is common to all experiences related to one object. By way of abstraction, we can make out what different objects have in common, and we can categorize them together into concepts. To build conceptual knowledge in this way is one of the core functions of abstraction (Burgoon et al., 2013; Yee, 2019). Concepts, in turn, are beneficial for many cognitive processes. First, they help to organize and structure the perceived environment (see also Barsalou et al., 2018). Second, they can help to reduce cognitive effort. When sorting similar things into the same category, the cognitive system can deal with the abstraction or concept, rather than with each individual instance. Third, concepts support generalization (Burgoon et al., 2013; Yee, 2019). For example, think of an agent who encounters a new object. Once the agent recognizes that this object belongs to a known category, the agent can form a hypothesis on how to interact with the new object. This means that abstraction is the process by which we build concepts, which in turn enable higher-level abstraction and generalization. The view that concepts are essentially abstracted knowledge also has implications for category structure. For example, we can think of the concept furniture to be more abstracted than the concept chair, which is in turn more abstracted than the concept office chair. The cognitive implications of these hierarchical, taxonomic relationships between concepts were investigated in early work in psychology by Rosch (1978) and Rosch et al. (1976). We can also think of the (abstract) concepts of freedom,friendship, or emotion as being more abstracted than the previous (concrete concept) examples. This is because they lack grounding in sensory-motor experience. These different kinds of abstracted knowledge can be characterized by the dimensions of specificity 2 , that is, the relationship between furniture,chair, and office chair, and concreteness 3 , that is, the qualitative difference between concrete concepts like chair and abstract concepts like freedom (Bolognesi et al., 2020). Recently, researchers began to believe that abstractness reflects a continuous dimension rather than a dichotomy between concrete and abstract concepts, a view that has dominated research on 1 In the following, I will use the term object only, but the claims also hold for other types of concepts, such as actions, events, and ideas. 2 The dimension of specificity relates to how specific or generic a concept is. Subordinate concepts are more specific and superordinate concepts are more generic. 3 The dimension of concreteness, or abstractness, relates to how concrete or abstract, in the sense of not directly perceivable, concepts are.
1 Motivation and outline 3 concepts in the previous decades (see also Banks et al., 2023; Barsalou et al., 2018; Burgoon et al., 2013; Yee, 2019). Through the lens of the conceptual hierarchy of more and less abstracted concepts, we can investigate the abstraction process involved. This process of conceptual abstraction is the key topic of this dissertation. Abstraction and language have a productive relationship. On the one hand, language is inherently abstract (Lupyan & Winter, 2018). In Linguistics, language is studied on different levels of abstraction, ranging from phonology, over morphology, up to syntax, pragmatics, and discourse. On the other hand, language and abstraction, especially in the sense of conceptual knowledge, are intertwined. Language is thought to play a role in forming concepts in at least three different ways. First, there is empirical evidence that language is useful for categorizing known concepts (e.g., Lupyan, 2005; Rissman & Lupyan, 2024). One example for this is the Russian blues study, which served as a motivating example at the beginning of this chapter (Winawer et al., 2007). Second, evidence suggests that language plays a role in forming new concepts. Specifically, it was demonstrated that language helps in learning (and teaching) new concepts (e.g., Arunachalam & Waxman, 2010; Callanan, 1985; Nazzi & Gopnik, 2001; Sloutsky & Deng, 2019; Sumers et al., 2023). In cognitive development, the acquisition of language and concepts goes hand in hand (Arunachalam & Waxman, 2010). Not only does the lexicalization of concepts help to collect information about a concept under one label, but language is also a source of acquiring information about concepts in the first place, for example, in reading or verbal teaching (Sloutsky & Deng, 2019; Waxman & Markow, 1995). The role of language seems to be larger the more abstracted (that is, generic and/or abstract) a concept is (e.g., Dobler et al., 2024; Gentner & Boroditsky, 2001; M. Lewis et al., 2021; Yee, 2019). For example, children learn superordinate concepts later than basic-level concepts, and they require language more to learn them (M. S. Horton & Markman, 1980; Mervis & Crisafi, 1982). This is related to the growing consensus that language is needed to explain the grounding of abstract concepts, which is believed to happen mainly via linguistic and social experience in interaction (e.g., Banks et al., 2023; Borghi & Binkofski, 2014; Borghi et al., 2017). Third, language was found to support successful generalization. For example, Suffill et al. (2022) related children’s knowledge of superordinate words to their inductive reasoning abilities. Taken together, these findings point to a role of language, and more specifically, interactive communication, in conceptual abstraction. Given that pragmatics plays a role when considering communicative interactions, I additionally consider the role of pragmatics in conceptual abstraction. Pragmatics is the study of meaning in context that goes beyond what was actually said, that is, what is semantically encoded in an utterance. In interaction, speakers typically do not use the most specific descriptions possible, but rather rely on the context of the communicative interaction and on the listener’s ability to infer the intended meaning. Pragmatic mechanisms play a role in communicating concepts on different time scales, ranging from language use to language acquisition and language evolution. On a short time scale, pragmatics plays a role in formulating and interpreting utterances in concrete communicative situations and contexts during language use (e.g., Graf et al., 2016). On a larger time scale, pragmatics is thought to help children acquire words and concepts in development (e.g., Bohn & Frank, 2019). And on an even larger time scale, pragmatics shapes an emerging language
10 2 Theoretical and empirical background similarity) and they share this property with basic-level categories (Murphy & Wisniewski, 1989; Rosch et al., 1976). Despite the importance of superordinate terms having sometimes been overlooked in prior research (Greene & Rohan, 2025), understanding the comprehension and production of superordinate terms is essential for understanding conceptual abstraction. Contrary to Greene and Rohan (2025), I take the fact that superordinate terms are acquired last by children (Mervis & Crisafi, 1982) not to mean that they are unimportant, but instead that they are challenging to acquire. Moreover, empirical evidence questions the primacy of the basic-level advantage (Macé et al., 2009; Murphy & Wisniewski, 1989; Poncet & Fabre-Thorpe, 2014). For example, it was shown that when categorizing objects in visual scenes, the basic-level effect decreases (Murphy & Wisniewski, 1989), or is even reversed, leading to a superordinate advantage (Macé et al., 2009; Poncet & Fabre-Thorpe, 2014). One explanation for these findings is that the basic-level advantage is a lexical-semantic effect and that the pathway of pure visual categorization is from coarse to fine, favoring the processing of superordinates (Macé et al., 2009; Poncet & Fabre-Thorpe, 2014). However, also in a task closer to the original object categorization task presented by Rosch et al. (1976) that features lexical-semantic information in the form of a label cue, the basic-level advantage decreases if superordinates are categorized in scenes (Murphy & Wisniewski, 1989). This suggests that the advantage of basic-level categories over superordinate categories in speed and accuracy decreases when superordinate categories are presented in their natural habitat, such as when referring to single objects in a scene or collections of objects (Wisniewski & Murphy, 1989, see also). These results demonstrate that the advantage of the basic over the superordinate level might not be as straightforward as suggested in the literature and that it might be worth looking into superordinate concepts more. 2.2 Labeling concepts in referential communication The studies investigating conceptual abstraction and the basic-level effect introduced in the previous section typically investigate the underlying concepts via reference. For example, in naming studies, participants are tasked to name objects depicted in images (Rosch et al., 1976). In categorization tasks, participants hear or read a label and then judge whether an object is an instance of the category that the label refers to. While categorization has also been investigated, for example, with a sorting task, where participants sort objects into categories without an explicit label (Rosch et al., 1976), much of the research on conceptual abstraction and categorization uses linguistic labels to investigate the mental and psychological constructs. Especially in ontogeny, it becomes clear that language and conceptual development are intricately linked. Children learn words and concepts simultaneously and in a reciprocal relationship where learning one facilitates learning the other (Arunachalam & Waxman, 2010). With the goal of this dissertation in mind, to investigate conceptual abstraction, I will examine the extensive literature on referential communication, which essentially involves labeling concepts.
2 Theoretical and empirical background 11 2.2.1 Studying referential communication in interaction The study of referential communication in Linguistics and pragmatics is based on two main assumptions that were introduced in the 1970s and 1980s. First, Grice (1975) introduced the idea that interlocutors in a conversation should be cooperative. Speakers are hypothesized to adhere to several maxims of conversation, which must be satisfied for communication to be successful, and for listeners to understand the implicated meaning in the speakers’ utterances. The quantity maxim states that a speaker’s utterance should be as informative as required for the listener to understand what is meant, but also not more informative. The maxim of quality urges speakers to stick to the truth and only utter what they believe to be true. The relation maxim is “[b]e relevant,” (Grice, 1975, p. 46), which means that speakers should only say what they deem relevant for the communicative purpose. Finally, the maxims of manner state that speakers should be precise and avoid lengthy or ambiguous expressions if possible (Grice, 1975). Second, H. H. Clark and Wilkes-Gibbs (1986) introduced the idea that referential communication is essentially collaborative. In referential communication, speakers and listeners talk about referents in the world, often using definite referring expressions such as the tall dog. The two partners in a conversation need to negotiate the referential process and establish a common perspective and mutual knowledge, which is called the common ground (H. H. Clark & Wilkes-Gibbs, 1986). These principles of interactive communication guide the speaker’s production and the listener’s interpretation of utterances. In modern experimental pragmatics, the formulation of these principles and observations has sparked research on reference resolution, that is, the comprehension of referring expressions, among other topics (Noveck & Reboul, 2008). For example, a vibrant field of research revolves around the question of how a shared perspective or common ground is established between interlocutors and how the (absence of) such mutual knowledge influences communication (e.g., Brennan & Clark, 1996; Brown-Schmidt & Hanna, 2011; H. H. Clark et al., 1983; Grodner & Sedivy, 2011; Katsos et al., 2023; Keysar et al., 2000; Richter et al., 2020). One central debate in experimental pragmatics is whether determining the speaker’s intended meaning requires effort on the listener’s side (Noveck & Reboul, 2008). The effort in a conversation can also lie with the speaker. Theories of audience design posit that speakers are willing to put in more effort in a conversation and cater their utterances to the listener’s needs (e.g., Ferreira, 2019; S. O. Yoon & Brown-Schmidt, 2019). Research investigating these communicative effects in interaction can shed light on these different accounts and theories. The most commonly used paradigm for studying referential communication in interaction is the reference game 1 . In a reference game, a speaker and a listener interact in a visual context. The speaker’s task is to send a message that helps the listener identify a particular target object from the visual context. The communication is successful if the listener interprets the message correctly and selects the correct target object. Reference games can be used to study either one-shot referring, where a speaker has only one chance to formulate a message, or repeated interactions, where speakers and listeners establish a referent over multiple turns. While the former is well-suited to investigate the pragmatic considerations at play during interpreting and producing a specific kind 1More information on the reference game paradigm will follow in Section 3.1.
12 2 Theoretical and empirical background of referring expression (e.g., Degen et al., 2020; Winters et al., 2018), the latter is particularly useful when investigating the process of conventionalization and entrainment (e.g., Barr & Kronmüller, 2006; Brennan & Clark, 1996; H. H. Clark & Wilkes-Gibbs, 1986; Hawkins et al., 2020). Referential communication provides an ideal testbed for studying the communicative principles formulated by Grice (1975) and H. H. Clark and Wilkes-Gibbs (1986) in a systematic and controlled way. 2.2.2 Conceptualization: The pragmatics of lexical choice Being interested in conceptual abstraction, in this dissertation I investigate a specific kind of referential communication where the referent is a concept on a particular level of abstraction. Referential communication, that is, how speakers and listeners communicate about objects in the world, has been investigated abundantly. However, even though the lexical choice of nouns is crucial for identifying an object and languages such as English need a noun to form a complete sentence, research on the choice of nominal referring expressions is relatively scarce (Graf et al., 2016). The literature mostly focused on modified nominal referring expressions, where nouns are modified with, for example, an adjective (e.g., Rubio-Fernandez et al., 2019,2022; Sedivy, 2003, 2005; Tourtouri et al., 2017, but see Graf et al., 2016; Gualdoni et al., 2022 for recent exceptions). Following Brennan and Clark (1996), the label that a speaker chooses to communicate an object with reflects which concept the speaker intends to communicate. For example, when calling an object a loafer, speakers conceptualize the object as a loafer and not as a shoe, or a piece of clothing (Brennan & Clark, 1996). Related to this, Barr and Kronmüller (2006) have argued that conversation, and specifically, referential communication, is one of the prime sites for category learning and category use. The authors argue that the choice of a specific label over another carries information. They call this the “’labeling as categorization’ hypothesis” (Barr & Kronmüller, 2006, p. 5). For example, by choosing the label mammal over the label whale, a speaker could highlight that whales are mammals and share attributes with mammals, such as needing to come up for air. In this dissertation, the questions of how speakers conceptualize a concept to communicate it and how listeners, in turn, understand it, and which information they gain from the labeling information are of central importance. From the literature in cognitive psychology and pragmatics, we know that there are several factors at play in the choice of a (communicatively) appropriate referring expression. In the classical view represented by Rosch et al. (1976), how an object is categorized and labeled depends on factors such as informativeness, lexical availability (i.e., the basic-level advantage), and perceptual salience (Brennan & Clark, 1996). These factors, particularly informativeness, are well-suited to a Gricean pragmatics account. When considering which label (loafer,shoe, or clothing) to choose for a particular object, these alternatives compete (Brennan & Clark, 1996). The basic-level effect literature predicts that speakers choose a basic-level term (shoe) over labels at the subordinate (loafer) or superordinate (clothing) level. According to pragmatic theory, the most appropriate referring expression depends on the context: If another type of shoe is present, the object should be referred to with the subordinate term loafer. This was also found in a recent experiment on nominal referring
2 Theoretical and empirical background 13 expressions conducted by Graf et al. (2016). They found that speakers preferred basic-level terms overall, that they used subordinate terms primarily when another object from the same basic-level category was present, and that speakers largely avoided producing superordinate terms (Graf et al., 2016). Under which circumstances a superordinate term becomes (pragmatically) relevant and appropriate in a conversation has long been overlooked. In this dissertation, I take inspiration from cognitive psychology, where it has been found that superordinate terms are used to refer to collections or classes of objects (Wisniewski & Murphy, 1989), to hone in on this question. This will allow me to draw a more complete picture of the cognitive and pragmatic factors at play when choosing a subordinate, basic-level, or superordinate referring expression. While the above-described factors do not depend on previous reference situations and the course of a conversation, in natural discourse, referring expressions are negotiated over repeated references in a collaborative setting (Barr & Kronmüller, 2006; Brennan & Clark, 1996). Barr and Kronmüller (2006, p. 7) state that “conversational settings shape speakers’ and listeners’ categorizations.” They discuss the importance of the fact that conversations are historical, multimodal, and collaborative for the process of categorization. In their view, categorizations are negotiated in communication, and the joint goal of the interlocutors is to establish a mutual perspective (Barr & Kronmüller, 2006). When bringing this view together with the idea that a referring expression reflects the conceptualization of an object (Brennan & Clark, 1996; E. V. Clark, 1997), the goal of a conversation can be formulated as aligning on a shared conceptualization. When this goal is achieved, speakers and listeners achieve communicative success and mutual understanding. In this dissertation, I am interested in how such shared conceptualizations are formed during interaction. In the case studies, we adopt the stance that categorizations are influenced by conversation and collaboration and investigate referential communication of concepts at different levels of abstraction in interactive experiments and multi-agent models. 2.3 The role of pragmatics in language and conceptual development The view that conversation and collaboration are important for categorization extends to language acquisition. Children acquire new words and their meanings during conversation and child-directed speech. Interdisciplinary research in cognitive science provides a particularly valuable framework to investigate linguistic and conceptual development jointly (Arunachalam & Waxman, 2010). Research on category and word learning established that children draw on both linguistic and observational information to learn the meaning of a novel word. This requires children to, on the one hand, identify the concept being referred to, and on the other hand, identify the word used to refer to the concept within a larger context. Finally, children need to establish a mapping between the word and the concept (Arunachalam & Waxman, 2010; Waxman, 1994). While conceptual development does not depend on language development, the two can interact and reinforce each other. In children’s typical development, concept and language acquisition are intricately linked, and words and concepts are learned simultaneously. For example, language may help to direct the learner’s attention to contrastive features and thus influence the categories being learned (Arunachalam & Waxman,
14 2 Theoretical and empirical background 2010; E. V. Clark, 1997; Waxman & Markow, 1995; Welder & Graham, 2006). One prominent feature of the modeling case studies presented in this dissertation is that we investigate the emergence of concepts and language jointly in interaction. Current proposals on language acquisition highlight the role of pragmatics therein (Bohn & Frank, 2019; E. V. Clark, 1997; Frank et al., 2009; Papafragou, 2002; Tomasello & Akhtar, 1995). Such socialpragmatic theories stand in contrast to association-based and constraint-based (Markman, 1994) theories of language acquisition (see Tomasello, 2000, for an argument contrasting social-pragmatic theories with associationand constraint-based theories). For example, Bohn and Frank (2019) argue that communication and pragmatics are central to language development in children as young as one year old. At first glance, this appears to be at odds with empirical evidence, which shows that children struggle with the prime test case of experimental pragmatics, namely the derivation of scalar implicatures (Horn, 2006; Noveck, 2001; Noveck & Reboul, 2008). The most used example for a scalar implicature is the implicature on the some/all scale. For adults, it has been reliably shown that they take a sentence like in (2.2) to mean that my niece ate some but not all of the cookies. My niece ate some of the cookies. (2.2) This means that adults pragmatically enrich the meaning of some, which is semantically compatible with all, to mean some but not all (e.g., Bott & Noveck, 2004; Horn, 2006; Noveck & Reboul, 2008; Papafragou & Musolino, 2003). The reasoning behind this meaning computation is spelled out as a combination of the reasoning about alternatives and the Gricean quantity maxim: If a speaker could have used the more informative term all, they would have done so. Therefore, the more informative all must be false in this context. This reasoning leads to the negation of the lexical alternative all in the interpretation of some. This mechanism can be generalized from scalar implicatures where lexical alternatives lie on the same scale, for example <some,all>, to conversational implicatures where the relevant alternatives can also be alternative referents in a context (Gotzner et al., 2020). Children, as opposed to adults, struggle to derive scalar implicatures (e.g., Noveck, 2001; Papafragou & Musolino, 2003), which has been taken as evidence that pragmatic development in children is gradual and comes much later (around age 7) than semantic development. How can these results be reconciled with the view that children at age one already use pragmatics to learn new words and concepts? Bohn and Frank (2019) review evidence for the view that pragmatics in the form of inference about the speaker’s goals already helps very young children to acquire language. They propose that even infants already understand the communicative function of language. Children expect speakers to use words intentionally, that is, to achieve a particular goal. For example, Vouloumanos et al. (2012) show that 12-month-old infants already attribute a communicative function to language. They let the infants watch a video where an actor failed to achieve their goal. Subsequently, an observer used unknown words vs. non-linguistic sounds (such as coughing). The infants assumed that the unknown words, but not the non-linguistic sounds, were instructions and predicted that the actor would achieve their goal in the subsequent trial (Vouloumanos et al., 2012). In line with this, E. V. Clark (1997) reviews evidence supporting the view that one-year-old children use pragmatic cues to learn new words. They focus on children’s learning of multiple words for the
2 Theoretical and empirical background 15 same object, such as dalmatian,dog, and animal (E. V. Clark, 1997). Learning multiple words that lie on the same taxonomic hierarchy is fundamental to this dissertation’s research on conceptual abstraction. E. V. Clark (1997) proposes that children can learn multiple perspectives, that is, many names for the same object, from the beginning of language acquisition. This view contradicts a one-perspective account that assumes that children have a bias for learning only one word per object. This relates to the mutual exclusivity bias, a learning bias that leads children to assume that a novel word refers to a novel object and not to a known object (Markman & Wachtel, 1988). In contrast to this, the evidence reviewed by E. V. Clark (1997) shows that as long as the words differ slightly in meaning and speakers give sufficient pragmatic cues, children have no difficulty learning multiple names for the same object. E. V. Clark (1997) conceptualizes multiple names for the same object as multiple conceptual perspectives a speaker can take and communicate. One main conclusion is that pragmatic cues are especially relevant for learning these multiple perspectives. The idea that children need sufficient cues to make use of pragmatic reasoning is in line with recent evidence for the early computation of conversational implicatures in young children. For example, Skordos and Papafragou (2016) and Gotzner et al. (2020) show that if lexical and contextual alternatives (such as all for the some-scalar-implicature or alternative objects in a context for exhaustivity implicatures) are made accessible and relevant, 4to 5-year-old children are able to derive implicatures. In summary, conceptual and word learning are intricately linked, and pragmatics plays a crucial role in language and conceptual development.
3 Experimental and modeling frameworks 3.1 Communication and reference games Referential communication has been investigated empirically with a communication task as early as the 1960s (Krauss & Glucksberg, 1969; Krauss & Weinheimer, 1964). Based on Wittgenstein (1959)’s idea that the meanings of words have their origins in their use in language and are determined in so-called language-games (orig. Sprachspiele), researchers have devised an experimental paradigm to investigate reference in interaction (Krauss & Glucksberg, 1969; Krauss & Weinheimer, 1964). 1 At the same time, the so-called signaling game was introduced as a game-theoretic model of how agents establish communicative conventions (D. K. Lewis, 1969). The Lewis signaling game can be formalized as follows. There are two players, a sender and a receiver. There is a world that can be in one state from the set of possible states 𝑆={𝑠1, 𝑠2, ..., 𝑠𝑛} . The sender knows the current state of the world, but the receiver does not. The speaker’s task is to choose a signal, or message, from a set of possible messages 𝑀={𝑚1, 𝑚2, ..., 𝑚𝑘} and use it to signal the world’s state to the receiver. The receiver’s task is to receive the signal and choose an action from a set of possible actions 𝐴={𝑎1, 𝑎2, ..., 𝑎𝑖} based on the signal. The sender thus maps states to messages as a function 𝑓:𝑆→𝑀 , and the receiver maps messages to actions as a function 𝑔:𝑀→𝐴 . There is exactly one correct action 𝑎𝑐 for each state 𝑠𝑛 . The joint payoff, or utility, of sender and receiver is 1 if the receiver selects the correct action 𝑎𝑐 and 0 if the receiver does not. Over repeated interactions, messages become conventionalized via coordination (D. K. Lewis, 1969). The experiment paradigm known as the communication game, the director-matcher task, or the reference game, follows the same general idea (see Figure 3.1). The speaker, or director, knows the true state of the world and has to communicate it to the listener, or matcher, who does not know the true state of the world (H. H. Clark & Wilkes-Gibbs, 1986). In H. H. Clark and Wilkes-Gibbs (1986), directors and matchers are allowed to make as many turns as they like to establish mutual knowledge. Reference games also exist in versions where the meaning of a signal is already established, and the use of this signal is investigated in a one-shot reference. Franke and Degen (2016) have proposed that reference games are instances of signaling games where the meaning is already conventionalized. But, in current research, the term reference game is employed in both use cases, that is, tasks that investigate language use with established conventions (see e.g., Achimova et al., 2022; Graf et al., 2016), and those that investigate the convergence on meanings during repeated interactions (see e.g., Boyce & Frank, 2023; Hawkins et al., 2020; Lazaridou et al., 2018; Ohmer, Duda, & Bruni, 2022). In this dissertation, reference games are used to study both language use with established meanings and the convergence on meanings during repeated interactions. 1 This communication game was later called the director-matcher task (H. H. Clark & Wilkes-Gibbs, 1986) and is nowadays often called the reference game in the context of referential communication.
18 3 Experimental and modeling frameworks Speaker Listener “blue” Figure 3.1: Visualization of the reference game. The speaker’s task is to communicate a target (here presented in a green frame) to a listener who has to select the correct object from the visual display. 3.2 Modeling language use with the Rational Speech Acts framework In pragmatics, language use can be modeled with the Rational Speech Acts (RSA) framework. The RSA modeling framework introduced in Frank and Goodman (2012) and Goodman and Stuhlmüller (2013) can be used to model language use as social cognition, that is, as inference over the speaker’s intentions and the listener’s likely interpretation of the speaker’s message. This and related approaches can be subsumed under the umbrella term probabilistic pragmatics (Franke & Jäger, 2016). They are based on game theory, such as the Lewis signaling game introduced above, and Bayesian cognitive modeling (Degen, 2023; Frank & Goodman, 2012; Franke, 2013; Franke & Jäger, 2016; Tenenbaum et al., 2011). Specifically, the RSA framework models pragmatic language use in interaction, where speakers reason about the listener’s likely interpretation of their utterance in language production, and listeners reason about the speaker’s intended meaning in language understanding. For more details on the formalization of this framework, given an example, see Box 1. Box 1. The Rational Speech Acts modeling paradigm Consider the reference situation in Figure 3.2 and let us assume that you only have a restricted number of words in your lexicon. You can choose one of the following expressions to refer to the objects in the display: {“blue”, “green”, “square”, “circle”} . What is the best expression for describing the target referent? Choosing either “blue” or “square” might seem like a good option to you because their literal semantic meaning matches the target referent. Still, you might not be fully satisfied with either of the two descriptions, as some ambiguity about the interpretation of the utterance remains. Specifically, “blue” could also refer to the circle, and “square” could also refer to the green square. As a pragmatic speaker, you would like to maximize the probability
3 Experimental and modeling frameworks 19 that your interlocutor selects the correct referent. This means that you need to reason about a listener’s likely interpretation of your possible utterances in this context. Assuming that your interlocutor is cooperative, they will, in turn, reason about your intentions when producing an utterance, which means they will try to infer the speaker’s intended referent in the context. Figure 3.2: An example for a reference game trial with one target displayed in the green frame and two distractors. This reasoning process can be modeled with Bayes’ rule (see Equation (3.1) ). A pragmatic listener tries to infer the state of the world 𝑤 given a speaker’s utterance 𝑢 by reasoning about the a priori likely world states 𝑃(𝑤) and the likelihood 𝑃𝑆(𝑢|𝑤) that a speaker chooses a particular utterance 𝑢to communicate a world state 𝑤: 𝑃𝐿(𝑤|𝑢) ∝ 𝑃𝑆(𝑢|𝑤)𝑃(𝑤).(3.1) The speaker is assumed to be rational (with an optional rationality parameter 𝛼 ), that is, to choose their utterance based on the utility 𝑈they expect to gain from a particular utterance: 𝑃𝑆(𝑢|𝑤) ∝ exp(𝛼𝑈(𝑢;𝑤)).(3.2) The expected utility 𝑈 is typically based on the Gricean assumption that speakers try to be informative. This means that their utility is high if their utterance enables the listener to select the correct state of the world 𝑤: 𝑈(𝑢;𝑤)=log 𝑃𝐿0(𝑤|𝑢).(3.3) 𝐿0 is the base case of the recursion of listeners reasoning about speakers who are in turn reasoning about listeners. Usually, the base listener 𝐿0 is modeled as a literal listener who interprets any utterance as their literal semantic meaning (Goodman and Frank, 2016, see also Degen, 2023 and Scontras et al., 2021 for an introduction to RSA models). What is the utterance with the highest utility in this context? A rational, pragmatic speaker is expected to choose the utterance “blue”. This can be modeled via the recursive inference as described above, but the problem also has an intuitively appealing solution: Two objects in the visual display are blue. However, if the speaker had wanted to refer to the blue circle, a more informative utterance would have been available, namely, “circle”. This form of pragmatic reasoning is similar to the reasoning about alternatives in scalar implicature derivation discussed above. The RSA framework is grounded in assumptions on general cognition, as well as pragmatic assumptions about meaning interpretation in context and interaction. It has been shown to fit the
26 3 Experimental and modeling frameworks 3.3.2 The emergent communication model used in this dissertation In this dissertation, I use a computational emergent communication paradigm to model the emergence of language and concepts as an interlinked and interactive process. I use the same basic model throughout case studies 2-4 presented in this dissertation. The models presented in the different case studies differ from the basic model only in ways that help to answer the respective research question of the case study. Using the same basic model also allows me to compare the models used in the case studies with each other and draw conclusions for the general research question of this dissertation. The base model is presented in Figure 3.3. Speaker neural network (GRU) Listener neural network (GRU) message [4,11,7,0] prediction and training 1 2 3 4 5 6 Figure 3.3: Base model for language emergence (case studies 2-4). The model consists of two agents, a speaker 𝑆 (1) and a listener 𝐿 (2). Both agents are implemented as a GRU (Cho et al., 2014). We choose the reference game paradigm because it is also used in the relevant experimental literature. As is common in the reference game paradigm, the speaker has to send a message to the listener that helps them to identify the intended referent. The messages (3) that our agents can send consist of discrete symbols from a vocabulary 𝑉 that consists of primitive discrete symbols, ranging from 0 to 𝑉 : "0", "1", "2", etc. (Lazaridou et al., 2018). Using a discrete communication channel has the advantage (over continuous communication) that agents cannot observe directly the inner states of their interlocutors. This makes discrete communication more human-like than continuous communication (where the message is represented by a continuous vector) as human communication is also discrete (Hockett, 1960; Lazaridou & Baroni, 2020). One additional advantage is that the combination of discrete symbols can result in compositional linguistic strategies (Lazaridou & Baroni, 2020). Speaker agents in our simulations can string up to 𝑀 symbols together, that is, they can produce messages with varying lengths up to a fixed maximum message length 𝑀 . The production and processing of these messages is handled by the GRUs. At each time step, a probability distribution over the next possible symbol from the vocabulary is computed based on the sequence so far, and the symbol with the highest probability is selected to resume the sequence. The symbol 0 is treated as an end-of-sequence (EOS) symbol and can end a message before 𝑀 is reached. In the reference game we employ, speakers and
3 Experimental and modeling frameworks 27 listeners communicate not about single targets, but rather about target concepts that consist of multiple objects. The speaker input (4) and the listener input (5) consist of 𝑔 target objects and 𝑔 distractor objects. We call 𝑔 the game size and set 𝑔=10 in our simulations. Speakers receive 10 target objects and 10 distractor objects in an ordered fashion, while listeners receive the same number of targets and distractors shuffled. Both speakers and listeners must learn the concepts in the world, that is, which target objects belong together, alongside learning to communicate about them. We use visualizations of objects of different shapes, colors, and sizes throughout the case studies. However, the agents receive symbolic, not visual, input data. We encode symbolic objects as tuples of 𝑛 attributes that can each take 𝑘 values. The speaker gets as input only the set of objects. The listener receives the set of objects and a message generated by the speaker. We use simple FF-NNs to embed the symbolic input data. The main reason why we use symbolic input data is that we want to control and manipulate different aspects of the input systematically. Our goal is to systematically construct input concepts at different levels of abstraction (dalmatian,dog, animal) and present them in systematically manipulated contexts. Constructing these inputs as symbolic inputs is a first step, but our work can also be extended to visual inputs. Another critical advantage of symbolic inputs is, however, that agents have been shown to focus on low-level visual features in past simulations. In these cases, agents have not learned the relevant concepts, but rather have converged on a communication protocol that, for example, communicates about two pixels instead of the entire image (e.g., Bouchacourt & Baroni, 2018; Lazaridou & Baroni, 2020). While this strategy is arguably very efficient (Lazaridou & Baroni, 2020), this is something that we do not want in our case studies because we are interested in studying how concepts are communicated. We can prevent the agents from communicating about low-level visual features instead of concepts by controlling all aspects of the input. This would not be possible (or at least much harder) with visual image data. We train (6) the agents in a RL paradigm. The primary teaching signal in RL is a reward for succeeding or failing a given task. In our model, the task is successful communication, and the reward is positive if listeners succeed in selecting the correct target objects. We can call this modeling framework ‘multi-agent’, as the agents treat each other as part of the environment and are not able to access the internal states of other agents (Lazaridou & Baroni, 2020). We use the Emergence of lanGuage in Games (EGG) toolkit (Kharitonov et al., 2019) for implementing the emergent communication models. The above-presented modeling paradigm allows us to study language and concept emergence jointly as an interactive process. While the agents do not know the target concepts, or which objects form a target concept, nor do their messages have meaning, the agents learn to communicate about target concepts via repeated interactions. At the beginning of training, the speaker agents send random messages in response to the data input, and the listener agents make random predictions about the objects in their input. During their joint training on communicative success, agents learn the concepts and develop a lexicon with mappings between concepts and messages that are commonly used to refer to them. Readers might ask why we use such small-scale models in times when Large Language Models (LLMs) are commonly available. Our choice is motivated by three main considerations. First, our main objective is to use computational modeling to gain insights and generate hypotheses about communication systems under controlled experimental conditions. The languages emerging in our
28 3 Experimental and modeling frameworks case studies are not intended to be full-fledged languages. Instead, we examine specific properties of these systems in simplified settings under carefully controlled conditions. By contrast, LLMs operate in more complex natural language, and are the subject of investigations themselves (see e.g., Holterman & van Deemter, 2023). While our ultimate aim is to develop explanations that apply broadly to communication systems, starting with controlled, small-scale models allows us to isolate relevant factors before addressing the full complexity of natural language. Second, from a theoretical viewpoint, our modeling framework closely parallels experimental paradigms used to study language evolution, such as the artificial language learning paradigm. This alignment facilitates comparisons with existing findings. Third, from a practical perspective, the small networks we use are sufficient for the input data and research questions we address. Using larger models would not yield additional insights but would require vastly more computational resources. LLMs require substantial amounts of compute time, not only during training, but also during inference (Bhardwaj et al., 2025), with significant implications for sustainability. Such resources should be employed only when necessary - a view consistent with the classic Occam’s razor approach to computational modeling, where the simplest model is preferred over the more complex model if other relevant parameters are equal. A good computational model should balance complexity and simplicity. This means that the model should be complex enough to capture the phenomenon of interest (in our case, the emergence of a communication system), and it should be simple enough to be interpretable (to a certain degree). Importantly, our analysis is conducted at Marr’s computational level (1982). We examine how the inputs and conditions we design change the observable and analyzable properties of an emergent communication system. We do not attempt to explain, at the algorithmic level, how the ANNs compute solutions. Our focus is on using computational modeling as a tool for understanding linguistic and cognitive phenomena. As argued by van Rooij et al. (2024), valuable insights in Cognitive Science can still be gained by employing more traditional AI methods for computational modeling in conjunction with critical theorizing.
4 Dissertation overview This dissertation investigates the role of language and pragmatics in conceptual abstraction - the process of forming and communicating concepts at different levels of a conceptual hierarchy. It builds on prior research in categorization, lexical choice in referential communication, that is, conceptualization, and the role of pragmatics and social interaction in language development and language evolution. In the previous chapters, I reviewed the literature on conceptual abstraction as well as the experimental and computational modeling frameworks that have been used to tackle questions on conceptual abstraction or will be used in the case studies presented in this dissertation. Five main conclusions from this review inform the design of the case studies: ▶ Cognitive economy and pragmatic informativity considerations play a role in conceptual abstraction. ▶ Referential communication involves conceptualizing and labeling concepts and therefore provides a controlled environment for testing questions about conceptual abstraction. ▶ Collaborative and Gricean pragmatic interaction plays a role in forming shared conceptualizations and conventionalized labels. ▶ Conceptual abstraction is shaped not only by environmental factors but also by communicative needs arising from interaction. ▶ Pragmatics plays an important role in language use, in language and conceptual development, and in language evolution. Accordingly, we examine how cognitive and pragmatic mechanisms influence both the actual use of language in concrete communicative situations and the emergence of structured linguistic and category systems. In the case studies, we employ experimental and computational modeling methods using the reference game paradigm to study referential communicative interaction. We investigate collaborative interaction and examine the extent to which pragmatic mechanisms are needed to establish effective and efficient shared conceptualizations and conventions. We go beyond investigating the basic-level effect and also focus on how informativity considerations make the suband superordinate levels of abstraction relevant for communication. We investigate language and conceptual evolution jointly to account for the assumption that conceptual abstraction is shaped by communicative needs and not only environmental factors. Finally, to gain a more holistic understanding of the role of language and pragmatics in conceptual abstraction, we investigate this role across different timescales, from language use to language evolution. The research presented in this dissertation is organized into four case studies, each taking a particular angle on the broader research question: “What is the role of language and pragmatics in conceptual abstraction?” Case study 1 investigates language use in interaction and, specifically, how people communicate about concepts at different levels of abstraction in a concept-level reference game. In case studies 2 to 4, we use agent-based computational modeling to investigate pragmatic factors that influence the structure of languages that emerged from interactive communication about
30 4 Dissertation overview concepts at different levels of abstraction. In case study 2, we are interested in the role of context in the emerging language and whether the availability of context leads to its use in an emerging communication system. Case study 3 extends the study of pragmatic context by studying recursive pragmatic reasoning about the intentions of the interlocutor modeled with the RSA paradigm. It brings together the study of language use and language emergence. Finally, in case study 4, we study abstraction as generalization via linguistic strategies. We ask how well agents generalize to novel levels of abstraction and which linguistic strategies help them achieve such generalization and abstraction. The following presents an overview of the specific research question, the methods used, and the timescales investigated in each of the four case studies. Box 3 below offers short definitions of the terminology that can also serve as a reference for later chapters. ▶ Case study 1: This case study investigates the cognitive and pragmatic factors at play when speakers and listeners communicate about concepts at different levels of abstraction in an interactive setting. We introduce the concept-level reference game, which is based on the classic reference game paradigm. This experiment investigates language use in concrete communicative situations, manipulating conceptual and contextual informativity. ▶ Case study 2: In this case study, we are interested in whether and how context availability shapes an emerging language. We introduce our agent-based model that is based on the emergent communication paradigm and the concept-level reference game. This framework tackles questions of language evolution by studying language emergence between agents. ▶ Case study 3: This is a comprehensive study divided into two experiments that examine the influence of different pragmatic mechanisms on efficient language use and language emergence. In Experiment 1, we use our framework to study the coevolution of category and linguistic systems and investigate the role of context-based pragmatics therein. In Experiment 2, we investigate the role of context-based and utility-based pragmatics modeled with the RSA framework on efficient language use. This study brings together emergent communication and RSA models in the two timescales of language evolution and language use. ▶ Case study 4: This case study explores the linguistic strategies that agents use when they need to communicate about concepts at novel levels of abstraction in a zero-shot generalization task. We use our agent-based emergent communication model and manipulate the training and test data splits to manipulate whether agents need to communicate about lower or higher levels of the conceptual hierarchy during language emergence and language use. Box 3. Operationalizations and modeling ingredients ▶ Communicative need - Arises from the relevant distinctions that need to be made in a language during interactive communication. ▶ Concept - Concepts can be regarded as abstracted knowledge, which they help to structure and organize. Concepts are operationalized as groups of multiple objects in experimental and modeling studies.
4 Dissertation overview 31 ▶ Concept-level reference game - A reference game where concepts need to be communicated instead of single objects. ▶ Conceptual hierarchy: experiment - Defined by inclusion relations, such as dalmatian ⊂ dog ⊂animal , where subordinate concepts lie on a lower level and superordinate concepts lie on a higher level of the hierarchy (see Figure 4.1). ▶ Conceptual hierarchy: modeling - Operationalized parallel to the experiment by the number of shared properties, or attributes, among members of the same concept. Specific concepts, where many attributes are shared, lie on a lower level and generic concepts, where only one attribute is shared, lie on a higher level of the conceptual hierarchy (see Figure 4.1). ▶ Context: experiment - The distractor objects in the reference game are defined by whether they lie on the same taxonomic hierarchy as the targets. For example, in a fine context, the distractors share the same basic-level category with the target objects (e.g., a dalmatian in the context of another type of dog).Inacoarse context, the distractors do not lie on the same taxonomic hierarchy (e.g., an animal in a context of a vehicle). ▶ Context: modeling - Operationalized parallel to the experiment by how many attributes the distractors share with the target concept. A context is fine if many attributes are shared between the distractor objects and the target concept. Conversely, a context is coarse if no attributes are shared between the distractor objects and the target concept. ▶Context-based pragmatics - The speaker agent’s access to a (shared) context. ▶ Efficiency - A concept from information theory (Shannon, 1948). Efficient communication is achieved when the message conveys the intended information with the least possible amount of resources, such as the length of the message or the size of the lexicon (e.g., Piantadosi et al., 2012). ▶ Emergent communication - A computational framework for modeling language evolution, see sections 3.3.1-3.3.2. ▶ Informativity - How much information can be gained from a label or message that describes a concept at a particular level of the conceptual hierarchy. Lower hierarchy levels are associated with higher informativitiy (in other words, we learn more about the concept) and higher hierarchy levels are associated with lower informativity (see Figure 4.1). ▶ Length cost - A cost applied to longer messages. Formalized as a loss pressure that penalizes longer messages by multiplying the symbol’s position in the message by a cost factor. ▶ Levels of abstraction - A term coined by Rosch et al. (1976) that is used to refer to different levels of a conceptual hierarchy. ▶Lexicon - A mapping between concepts and messages. ▶ Polysemy - A one-to-many mapping between messages and concepts where one message can refer to more than one concept. ▶ Reference game - An experimental paradigm for studying referential communication, see section 3.1. ▶ Rational Speech Acts (RSA) - A modeling paradigm for modeling referential communication based on Bayesian cognitive modeling, see section 3.2. ▶ Shared context - A context is considered to be shared between speakers and listeners when
32 4 Dissertation overview the attribute(s) that differ between the distractor objects in the context and the target concept are the same in the speaker’s and the listener’s input. ▶ Synonymy - A many-to-one mapping between messages and concepts where more than one message can be used to refer to the same concept. ▶ Utility - How useful a message is for communication in a particular situation or context. The utility of an utterance is influenced by the listener’s likely interpretation of an utterance (see Section 3.2). Other objectives, such as a cost for longer utterances, can be included in the utility function of an RSA model as well. ▶ Utility-based pragmatics - Operationalized by implementing RSA speakers, that is, speaker agents that maximize their utterances’ utility via an RSA-based production strategy considering the listener’s likely interpretation of each potential utterance. ▶ Zero-shot generalization - A paradigm for testing ML models on unseen categories, tasks, or conditions. In contrast to standard testing, the data does not come from the same distribution as the training data. low high informativity high low conceptual hierarchy Dalmatian Animal Dog superordinate basic subordinate few many generic specific attributes shared (within concept) Labels Experimental stimuli Modeling stimuli Figure 4.1: A comparison of how the conceptual hierarchies are operationalized in the experiment and modeling studies. Concepts at higher levels of the conceptual hierarchy are superordinate or generic concepts, which are associated with low informativity. Concepts at lower levels of the conceptual hierarchy are subordinate or specific concepts, which are associated with high informativity. In the following chapters, the case studies will be presented. Each case study is preceded by a high-level introduction to the subject. After each case study, I sum up the key findings and describe the implications for the broader research question regarding the role of language and pragmatics for conceptual abstraction. The final chapter of this dissertation is a discussion of the key contributions and the ramifications of each case study for the role of language and pragmatics in conceptual abstraction. I will also deliberate on the implications of the presented case studies for ongoing and future research in Cognitive Science, (Computational) Linguistics, and pragmatics, and finally draw a general conclusion about the here presented work.
5 Case study 1: The cognitive and pragmatic factors in referring to concepts at different levels of abstraction This chapter presents case study 1. It starts with a high-level introduction followed by the content of the publication: Kobrock, K., Uhlemann, C., & Gotzner, N. (2024). Superordinate referring expressions in abstraction: Introducing the concept-level reference game. Proceedings of the Annual Meeting of the Cognitive Science Society,46, 518–525. https://escholarship.org/uc/item/31n5d3p6 The chapter ends with a brief summary of the main contributions of the publication and a discussion of its implications for the broader research question of this dissertation. 5.1 High-level introduction During natural conversation, speakers face choices about which words to use to communicate what they intend. According to pragmatic theory and the classical Gricean maxims (Grice, 1975), speakers are cooperative and want to be understood by their interlocutor. This means that they try to communicate as much information as necessary for the listener to understand what the speaker means. Speakers also try to minimize their effort and communicate not more than necessary (Grice, 1975). These principles guide the selection of words during conversation and ensure that communication is successful. In this work, we look at a specific kind of choice speakers face when they search for the right word to convey a particular concept. We focus on nominal referring expressions that are used to communicate concepts at different levels of abstraction (Rosch et al., 1976). For example, speakers can use the subordinate term dalmatian, the basic-level term dog, or the superordinate term animal to refer to the same dog in the real world (see Figure 5.1). In this publication, we experimentally investigate the speaker’s choice of referring expression for concepts at different levels of abstraction. We pit against each other the speaker’s need for being informative with their need for low effort. Maximizing informativeness, a speaker would always use the most specific term that is possible to use in a given context. Minimizing effort, speakers would communicate terms that are easiest and fastest for them to produce. These are terms at the basic level, such as dog,car, and chair (Rosch et al., 1976). Our hypothesis is that the speaker’s preference for basic-level terms competes with the speaker’s desire to be informative. We construct reference situations in which speakers have to use a term at a certain level of abstraction to successfully communicate an intended referent. These are visualized in Figure 5.1. For example, if the speaker wants to refer to a dalmatian, but another type of dog is present, the speaker has to choose the subordinate referring expression dalmatian because the basic-level term dog would not be informative enough. To make also superordinate referring expressions such as animal informative in certain contexts, we implement situations where two kinds of animals need to be communicated
34 5 Case study 1: Referring to concepts at different levels of abstraction together. For example, if speakers want to communicate a parrot and a dog, they would need to use the superordinate referring expression animal in order to be both informative, and economical (expressions using conjunctions such as dog and parrot are possible but require more effort because they are longer). We look at both sides of a conversation, speakers and listeners, because we want to understand how listeners comprehend the referring expressions chosen by the speakers. By looking at the listener side of the conversation, we can determine the communicative success of a speaker’s utterance, and investigate whether listeners also have a preferred level of abstraction at which they can understand the intended target concepts better. If listeners process referring expressions at the basic level faster than of expressions at the subordinate or superordinate levels, this would indicate that listeners also have a preference for basic-level terms and that speakers might choose basic-level terms partly because they make comprehension easier and help them achieve communicative success. Conversely, if the basic-level is not processed faster than subor superordinate terms, then this means that listeners do not benefit from the speakers using basic-level terms. This could mean that the usage of basic-level terms is only advantageous for the speakers and that they use it mainly to minimize their own production effort. However, every time a speaker deviates from the basic level for the sake of communicative success, it would indicate that speakers are willing to produce expressions that are more effortful for them. One explanation for this behavior could be that speakers are trying to be cooperative. They know that a subor superordinate referring expression can enhance communicative success in certain reference situations, allowing the listeners to select the intended referents. Our experimental design allows us to investigate these factors in communicative interaction. basic-level advantage conceptual informativity contextual informativity superordinate Dalmatian Animal Dog basic subordinate Figure 5.1: Visualization of the levels of abstraction and the cognitive and communicative pressures for the conceptual hierarchy dalmatian ⊂dog ⊂animal . The image of a dalmatian is taken from the BOSS image database licensed CC BY 4.0 (Brodeur et al., 2014). The other stimuli are stimuli employed in the study and reused from Graf et al. (2016).
5 Case study 1: Referring to concepts at different levels of abstraction 35 5.2 Abstract We study referential communication about concepts at different levels of abstraction in an interactive concept-level reference game. To better understand processes of abstraction, we investigate superordinate referring expressions (animal). Previous work identified two main factors that influence speakers’ choice of referring expressions for concepts: the immediate context and the basic-level effect, i.e. a preference for basic-level terms such as dog. Here we introduce a new concept-level reference game that allows us to study differences in the basic-level effect between comprehension and production and to elicit superordinate referring expressions experimentally. We find that superordinate referring expressions become relevant for groups of objects. Further, we reproduce the basic-level effect in production but not in comprehension. In conclusion, even though basic-level terms are most readily accessible, speakers tailor their expressions to the context, allowing the listener to identify the target concept. Keywords: reference game, concepts, categorization, superordinate level, abstraction 5.3 Introduction Concepts allow us to make sense of the world. They help us to structure and organize knowledge, and to generalize from one instance to a class of objects that share similar properties through a process that is commonly called “abstraction” (Rosch, 1978; Yee, 2019). We use referring expressions at different levels of abstraction, ranging from subordinate terms like dalmatian to superordinate ones like animal, to communicate about concepts at different levels of abstraction. Previous work suggests two main factors that influence the choice of referring expressions (REs) people use to refer to concepts at different levels of abstraction. On the one hand, Rosch et al. (1976) famously found that basic categories are special because these are “the most inclusive categories for which a concrete image of the category as a whole can be formed” (Rosch et al., 1976). It has been shown that children acquire basic-level terms like dog first (J. M. Clark & Johnson, 1994; Mervis & Crisafi, 1982) and that objects can be categorized faster at the basic level than at the subor superordinate levels (Murphy & Smith, 1982). On the other hand, the Gricean maxim of quantity predicts that speakers provide as much information as required for the listener to identify a target in a given context and not more (Grice, 1975). This means that speakers should tailor their utterances to the communicative situation at hand, considering both the concept they would like to communicate and the context of their utterance. It has also been shown empirically that context plays a role in the selection of REs in referential situations (see for example Hawkins et al., 2018,2020; Konopka & Brown-Schmidt, 2014; Sedivy, 2005). Our goal is to pit the two factors directly against each other. We use a reference game similar to Graf et al. (2016), where a speaker describes an object and a listener needs to identify this object from a set of distractors. Graf et al. (2016) showed that while speakers tailor their utterances to the context, they also prefer basic-level expressions (e.g, dog) overall.
42 5 Case study 1: Referring to concepts at different levels of abstraction 98.93% and 90.29%, respectively. We find substantial evidence for the differences between both the basic and fine conceptual context (M=1.39, CrI=[0.34, 2.34], pd=99.08%, ROPE=[-0.18, 0.18, 0% in ROPE]) and between the coarse and fine conceptual context (M=2.3, CrI=[0.33, 4.35], pd=99.11%, ROPE=[-0.18, 0.18], 0% in ROPE). 6 The effect of conceptual context on the reference level is further supported by a Bayes Factor of 13.5 in favor of the model that includes conceptual context as a predictor against a null model, providing strong evidence for the effect of conceptual context on the choice of the reference level.7 5.5.4 Response times For the response time analyses, we were interested in response times of trials in which the listener selected the correct target objects, i.e. communication was successful, and in which the speaker chose an utterance on the appropriate level of abstraction. The exclusion of data on inappropriate levels led to a data loss of 11.12%. Speaker response times Our second hypothesis was that speakers choose utterances on the basic level more quickly than utterances on subor superordinate levels. This hypothesis was motivated by the basic-level effect found in the literature (Rosch et al., 1976; Tanaka & Taylor, 1991). We thus predicted that speaker response times would be shorter when the produced utterance was on the basic level compared to the other two levels. As preregistered, trials with response times over 2.5 standard deviations above the mean (cut-off: 22,360 ms) were excluded, leading to the exclusion of 2.61% of the remaining data. Figure 5.5 shows the data means and bootstrapped confidence intervals of the cleaned data. Indeed, responses on the basic level are shorter than response times on the other two levels. We ran a Bayesian model with a lognormal link function, predicting speaker response times by conceptual context and including group-level effects for the participant pair and item category. We specified weakly informative priors to enhance model fit. We had to deviate from one prior specified in the preregistration: Prior predictive checks showed that we overestimated the effect size, and the prior we preregistered for the population-level effect was too wide. We thus changed the prior’s standard deviation from 2 to 0.5 to enhance model convergence and fit. 8 The model results 6 At the suggestion of our reviewers, we also ran a model on the data without excluding unsuccessful trials. This model provides similar evidence to the original model: basic-fine: M=1.63, CrI=[0.35, 2.81], pd=98.65%, ROPE=[-0.18, 0.18], 0% in ROPE; coarse-fine: M=2.58, CrI=[0.44, 4.79], pd=99.16%, ROPE=[-0.18, 0.18], 0% in ROPE. 7 The BF models ran for 20,000 iterations with a warm-up period of 2,000 as recommended for BF estimation (Nicenboim et al., 2023). 8 The priors were specified as follows: intercept prior: normal(8.65, 0.5), population-level effects slop prior: normal(0, 0.5), group-level effects standard deviation prior: normal(0, 0.1), group-level effects correlation prior: lkj(2), sigma prior: normal(0, 0.5).
5 Case study 1: Referring to concepts at different levels of abstraction 43 Figure 5.5: Speaker response times with bootstrapped 95% CIs. confirm that speakers respond substantially faster on the basic level than on the other two levels of reference (M=-0.22, CrI=[-0.36, -0.10], pd=99.86%, ROPE=[-0.01, 0.01], 0% in ROPE).910 Listener response times Our third hypothesis was that we would also find differences between levels in the listener response times, showing that either a) listeners select the correct targets more quickly when a basic-level term was produced or b) listeners select the correct targets more quickly when a superordinate term was produced. These hypotheses were based on the basic-level effects literature (Rosch et al., 1976) and more recent studies debating the basic-level effects in certain tasks (Macé et al., 2009). The data in Figure 5.6 shows that we do not find either of the expected patterns. Figure 5.6: Listener response times with bootstrapped 95% CIs. 9This posterior difference was contrast-coded as preregistered: basic vs. (fine + coarse)/2. 10 At the suggestion of our reviewers, we also ran a model including length as a predictor. This model provides smaller evidence for the difference of the basic level compared to the other two: M=-0.10, CrI=[-0.19, 0.00], pd=98.07%, ROPE=[-0.01, 0.01], 0.62% in ROPE. This is mostly driven by the difference between basic and coarse being not as pronounced as the difference between basic and fine.
44 5 Case study 1: Referring to concepts at different levels of abstraction We ran a Bayesian model with the same model specifications and priors as the speaker response time model. 11 We employed the same data exclusion criteria as for the speaker response time model with one exception: We decided to exclude the highest response time data point and defined the cut-off of 2.5 standard deviations above the mean after this exclusion to get a more sensible cut-off (17,076 ms). The model predictions do not offer enough reason to believe that either of the two proposed patterns is at play: The difference between the basic level and the other two levels was estimated at M=-0.04 with a CrI of [-0.12, 0.04] (pd=84.96%, ROPE=[-0.01, 0.01], 12.28% in ROPE). And the difference between the coarse level and the other two levels was estimated at M=0.04 with a CrI of [-0.07, 0.14] (pd=79.07%, ROPE=[-0.01, 0.01], 11.81% in ROPE).12 5.6 Discussion and conclusion We studied the expressions speakers use to refer to concepts at different levels of abstraction in a new interactive concept-level reference game. We have shown that the level of abstraction of the utterances speakers choose to communicate a certain concept mainly depends on the concept and context in question. Speakers tailor their utterances to the context, producing subordinate terms in fine conceptual contexts, basic-level terms in basic conceptual contexts and superordinate terms in coarse conceptual contexts. In fact, we found that if speakers produce a RE on the expected level of reference, this reduces errors in the target selection by the listeners, or, in other words, it increases communicative success by about 14%. We also reproduced the basic-level effect on the production side, i.e. speakers are faster in producing a basic-level term than in producing a term on the other two levels. Interestingly, even though speakers show a basic-level advantage, they still tailor the utterances to fit the context and make it easier for the listener to identify the correct targets. We did not find evidence for an advantage of basic-level processing on the comprehension side, i.e. listeners are equally quick in selecting the targets regardless of the level of abstraction of the utterance they receive. A possible reason for this is that the basic-level advantage is mostly driven by accessibility in production. In comprehension, on the other hand, listeners might be as quick in categorizing objects at other levels than at the basic level. Despite the basic-level effect, speakers use suband superordinate terms frequently in natural conversation. In the case of subordinate terms, this is usually determined by the context. When, for example, a dalmatian is the target and a greyhound is the distractor, the basic-level term dog does not sufficiently discriminate the target from the distractor. It has been shown that context warrants the use of a more costly, i.e. usually longer and less frequent, subordinate term (Graf et al., 2016). In the case of superordinate terms, however, the context does not sufficiently explain why these terms might be used because participants could always use the basic-level term to refer to single objects even in a coarse conceptual context. Our concept-level reference game shows that superordinate 11 The only exception was the prior on the intercept that was preregistered to depend on the data distribution: normal(8.15, 0.38). 12 These posterior differences were contrast-coded as preregistered: basic vs. (fine + coarse)/2 and coarse vs. (basic + fine)/2.
5 Case study 1: Referring to concepts at different levels of abstraction 45 terms become relevant when dealing with multiple target objects, or when communicating the idea of a more generic class. The longer response times we see when speakers use a superordinate compared to a basic-level term might be an indicator of a process of abstraction that speakers undergo when trying to find which superordinate class the two targets have in common and retrieving the respective superordinate RE. This could be a good starting point for further research on abstraction. One limitation of our current study setup is that so far, we only investigate REs in a context that is very close to the target concept, i.e. that includes a distractor from the same basic category in the fine conceptual context or from the same superordinate category in the basic conceptual context. This means that we cannot account for overinformative REs (see for example Degen et al., 2020) because the context makes a certain level of reference necessary for disambiguation. Future studies can extend our setup and include context conditions that make a certain level of reference only sufficient for discrimination, by using wider contexts that, for example, only include unrelated distractors. Such a manipulation would allow the investigation of overand underspecification in the concept-level reference game. However, even in our current set-up, we do see some utterance choices that are at odds with our predicted level of reference for each conceptual context. For example, in the fine conceptual context, basic-level expressions are produced almost 20% of the time. A closer look at these productions reveals that speakers either underspecify, i.e. produce dog for dalmatian, or they use a modified basic-level expression, i.e. produce spotted dog. The high proportion of these mentions provides further evidence for a strong basic-level effect on the speaker side. The response time results in which we find a basic-level effect only for production, but not for comprehension, lead to an interesting observation: Speakers tailor their utterances to the conceptual context even when it results in higher processing costs for them. On the comprehension side, however, we do not find higher processing costs for subor superordinate terms. This could suggest that speakers are willing to bear a higher cost because they know that it would make identification of the target objects easier for the listener. This phenomenon has been discussed in the literature as audience design (see for example Gann & Barr, 2014; W. S. Horton & Gerrig, 2002). We should note, however, that the comprehension response times were logged when listeners had clicked on both targets. Thus, our measure is rather offline, and we cannot completely rule out that there are more immediate differences across levels in comprehension that might be revealed by more sensitive measures. If however the differences we observed for production and comprehension are not just due to such methodological aspects, this could indicate that the basic-level advantage is related to lexical accessibility and not categorization itself. On the listener’s side, basic level categories may not have a privileged representation. In conclusion, the concept-level reference game allows us to test hypotheses on the use of superordinate REs and abstraction. While we see differences in response times between the basic and superordinate levels on the production side, we do not see the same differences on the comprehension side. This opens up exciting possibilities for future research on audience design and costs
46 5 Case study 1: Referring to concepts at different levels of abstraction associated with abstraction in production and comprehension. Here, we showed that superordinate REs become relevant when a speaker needs to describe more than one target object. Acknowledgments We would like to thank Ilva Hovemann for help with programming the experiment in Labvanced, Berit Reise for help with programming the norming study in pcIBEX, and Elli Tourtouri for helpful discussions on the experiment design. We also thank three anonymous reviewers for their helpful comments and feedback. Kristina Kobrock is supported by the DFG-funded Research Training Group “Computational Cognition” (DFG-GRK 2340). Author Contributions: Kristina Kobrock: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Writing - Original Draft, Visualization. Charlotte Uhlemann: Conceptualization, Methodology, Software, Writing - Review & Editing. Nicole Gotzner: Conceptualization, Methodology, Writing - Review & Editing, Supervision.
5 Case study 1: Referring to concepts at different levels of abstraction 47 5.7 Summary of the key contributions Weintroducedanovelexperimentalparadigmtoinvestigatehowspeakersandlistenerscommunicate about concepts at different levels of abstraction. We built on the classic reference game paradigm, which is commonly used to study referring expressions, and extended it to concept-level nominal reference. Our findings contribute to the literature on cognitive mechanisms, such as the basic-level effect, as well as pragmatic factors in producing and processing referring expressions, including context and informativity considerations. First, we replicated the well-known basic-level effect in production but not in comprehension. Second, we found evidence that superordinate referring expressions are used when relevant, that is, when they are conceptually informative. Third, when considering the tradeoff between the basic-level advantage and informativity considerations, we found that the pragmatic factors of conceptual and contextual informativity determine the speakers’ choice of referring expressions, even if they go against the basic-level effect. The speakers’ prevalent strategy of using referring expressions at the appropriate level of reference, licensed by the conceptual level and context, leads to higher communicative success than when speakers use referring expressions at an inappropriate level of reference. 5.8 Implications for the broader research question We addressed open questions in the empirical literature on the role of pragmatics in nominal reference (e.g. Degen et al., 2020; Graf et al., 2016). Our study goes beyond previous work by investigating reference to target concepts instead of single target objects. Moreover, as far as we know, we are the first to study the basic-level effect in interaction. This opens up the possibility to study how speakers communicate target concepts at the sub-, superordinate, and basic levels, and how listeners comprehend such target concepts. We found that speakers trade off pragmatic factors, specifically the desire to be informative, with cognitive factors, which relate to the ease of retrieving basic-level referring expressions. In our results, we see that speakers use a basic-level term where possible, but they also deviate from the basic-level production strategy and use terms at the suband superordinate levels of abstraction when these are more appropriate. If the target concept is a superordinate concept, then the conceptual informativity makes the superordinate term more informative than the basic-level term. If the target concept is subordinate and presented with a distractor object from the same basic-level category, then the contextual informativity makes the subordinate term more informative than the basic-level term. We found that speakers take these conceptual and contextual informativity considerations into account when referring to concepts. Speakers additionally show a basic-level advantage in response times which means that they tend to produce basic-level terms faster than terms at the subor superordinate levels. This corroborates the well-known basic-level effect when naming concepts in an interactive setting in production (Jolicoeur et al., 1984; Rosch et al., 1976). Interestingly, a potential strategy for the speakers in the superordinate condition would be to communicate both basic-level objects that comprise the target concept in a conjunction. However, the experimental data suggest that the superordinate term
48 5 Case study 1: Referring to concepts at different levels of abstraction is used more consistently and more successfully (mentions at the appropriate level of reference increase communicative success by 14%). This is in line with predictions of cognitive economy: The superordinate term is usually shorter than mentioning two basic-level terms. Thus, if there is no additional cost associated with retrieving the superordinate term, it should be cognitively more economical to produce it. In comprehension, we did not find any differences in response times when listeners interpreted appropriate referring expressions. Listeners seem to be equally able to identify the target concepts independent of their level of abstraction. This could mean that listeners do not face an abstraction effort when identifying superordinate concepts or a discrimination effort when discriminating subordinate concepts in contexts with close distractors during interaction. Together, these findings suggest that speakers might use referring expressions at levels other than the basic level primarily to achieve communicative success in interaction when there is a need to be more informative. In conclusion, we showed how cognitive mechanisms and pragmatic factors are pitted against each other in interaction. Our novel experimental paradigm allowed us to tackle the question of how concepts at different levels of abstraction are communicated and understood. In the following studies, we take this paradigm to the next step and use it to model how languages evolve during interactive communication about concepts.
6 Case study 2: The role of context availability in emergent communication about concepts This chapter presents case study 2. The chapter starts with a high-level introduction followed by the content of the publication: Kobrock, K., Ohmer, X., Bruni, E., & Gotzner, N. (2024). Context Shapes Emergent Communication about Concepts at Different Levels of Abstraction. In N. Calzolari, M. - Y. Kan, V. Hoste, A. Lenci, S. Sakti, & N. Xue (Eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 3831–3848). ELRA and ICCL. https://aclanthology.org/2024.lrec-main.339 The chapter ends with a brief summary of the main contributions of the publication and a discussion of its implications for the broader research question of this dissertation. 6.1 High-level introduction Inthepreviouscasestudy,weexaminedwhichreferringexpressionsspeakerschoosetocommunicate concepts in interactive communication and how listeners comprehend such referring expressions. The goal of the following study is to investigate communication about concepts at different levels of abstraction in a computational model of language evolution. This allows us to generalize our research findings from language use in specific communicative situations, as investigated in case study 1, to the question of how a language evolves to encode concepts at different levels of abstraction. The fact that we can refer to the same object with different names that lie on the same taxonomic hierarchy is puzzling to researchers studying language development and language evolution. From the perspective of language development, children typically expect one-to-one mappings between objects and words (Markman & Wachtel, 1988). From the perspective of language evolution, it remains an open research question which pressures influence the evolution of multiple names for the same object. The goal of case study 2 is to address this question. To achieve this aim, we implement a computational model of language evolution that enables us to test several hypotheses regarding how a language evolves to facilitate reference to different levels of abstraction. The key components of our model are: First, we model communication in interaction. We train artificial agents, one with the speaker role and one with the listener role, on the same communicative task as was used in case study 1, the concept-level reference game. Second, following our findings from case study 1, we model conceptual and contextual informativity. We use a dataset consisting of hierarchical concepts for the communicative task. We formalize concepts as groups of objects that share a certain number of attributes. We hypothesize that this conceptual hierarchy should allow a language to emerge that encodes concepts at different levels of abstraction. We implement contextual informativity by adding distractor objects. The context created by these distractor objects is characterized by the number of attributes shared between target and distractor objects. If many attributes are shared between target and distractor objects, then the context requires fine distinctions and a specific description of
50 6 Case study 2: The role of context availability in emergent communication about concepts the target concept. Conversely, if few attributes are shared between target and distractor objects, coarse distinctions and more generic descriptions of the target concept are sufficient for effective communication. In case study 2, our first aim is to introduce this computational model and the dataset, manipulating conceptual and contextual informativity. Our second aim is to investigate the role of the context and, specifically, whether the availability of context shapes an evolving language. We achieve this aim by implementing two experimental conditions that manipulate whether the speaker agent has access to context information or not.
6 Case study 2: The role of context availability in emergent communication about concepts 51 6.2 Abstract We study the communication of concepts at different levels of abstraction and in different contexts in an agent-based, interactive reference game. While playing a concept-level reference game, the neural network agents develop a communication system from scratch. We use a novel symbolic dataset that disentangles concept type (ranging from specific to generic) and context (ranging from fine to coarse) to study the influence of these factors on the emerging language. We compare two game scenarios: one in which speaker agents have access to context information (context-aware) and one in which the speaker agents do not have access to context information (context-unaware). First, we find that the agents learn higher-level concepts from the object inputs alone. Second, an analysis of the emergent communication system shows that only context-aware agents learn to communicate efficiently by adapting their messages to the context conditions and relying on context for unambiguous reference. Crucially, this behavior is not explicitly incentivized by the game, but efficient communication emerges and is driven by the availability of context alone. The emerging language we observe is reminiscent of evolutionary pressures on human languages and highlights the pivotal role of context in a communication system. Keywords: emergent communication, concepts, context 6.3 Introduction Referring to things in the world is crucial to effective communication. When choosing a referring expression, speakers recur to what they know about the referent’s underlying concept and choose to communicate the concept at a level of abstraction that fits well with their communicative intentions. For example, a reference to the object in Figure 6.1 can be made at various levels of abstraction, ranging from the more specific concept ‘watermelon’ to the more generic concept ‘food’. When communicating a more generic concept, speakers and listeners need to abstract away from properties of the individual objects and focus on what all objects belonging to a concept have in common. By choosing the utterance ‘fruit’, for example, a speaker abstracts away from irrelevant properties (the size, color etc. of the specific object) and stresses the properties that watermelons share with other items belonging to the concept ‘fruit’, for example that they are edible. “watermelon” “melon” “fruit” “food” generic specific Figure 6.1: Example referring expressions at different levels of abstraction.
58 6 Case study 2: The role of context availability in emergent communication about concepts when the listener correctly identifies most objects per game. First, we observe very high training and validation accuracies for all game settings and datasets (mean training and validation accuracies across runs > 0.96 for all datasets and both settings). 4 Mean test accuracies across runs on concepts that the agents have never encountered during training are 0.89 (SD=0.07) for context-unaware and 0.87 (SD=0.11) for context-aware agents. This suggests that the agents learn to successfully communicate about concepts on various levels of abstraction and in various context conditions. 5 To get a better understanding of the agents’ strategies and where communication is especially (un)successful, we performed an additional analysis of the errors, i.e. those cases where the listener agents predict some of the labels wrongly. 6 We find that most errors occur when targets and distractors share many attributes, making it more likely that they are confused with each other. In other words, most mistakes happen in the fine context conditions.7 6.6.2 Qualitative communication analysis Second, we use a qualitative analysis of the messages to see whether we can observe differences between context-unaware and context-aware settings. Tables 6.2 and 6.3 show the results of our qualitative analysis on the D(4,4) dataset for context-unaware and context-aware, respectively. We report all unique messages for a randomly chosen specific concept ( [0,0,0,3] , all attributes fixed) and each context condition. Context-unaware agents tend to use the same messages in all context conditions (in this case, “ [11,1,11,14,0] ” is used consistently across contexts). On the other hand, context-aware agents use a larger set of messages (four unique messages over all games), and they tend to vary the messages more depending on context. In coarser contexts, the set of messages used to describe the target concept is larger than in the finest context, where the best strategy is to communicate the most specific concept. We observe the same pattern for other randomly selected concepts across different datasets.8 Table 6.2: Context-unaware: Unique messages used to refer to a randomly picked specific concept in the D(4,4) dataset over different context conditions. Object Context(# SharedAttributes) UniqueMessages [0,0,0,3] 0 “[11,1,11,14,0]” 1 “[11,1,11,14,0]” 2 “[11,1,11,14,0]” 3 “[11,1,11,14,0]” 4 It is important to note that achieving such high scores is intentional. Only with a high success score does the rest of the evaluation become meaningful. This ensures that the language we analyze can be assumed to effectively communicate what is intended in the referential game. 5Detailed accuracy scores can be inspected in Table 6.4 in Appendix A. 6Plots for these analyses can be inspected in Appendix B. 7Additional plots of the distribution of false positive and false negative errors can be found in Appendix B.2 and B.3. 8More examples are given in Appendix Cto show that these are not cherry-picked.
6 Case study 2: The role of context availability in emergent communication about concepts 59 Table 6.3: Context-aware: Unique messages used to refer to a randomly picked specific concept in the D(4,4) dataset over different context conditions. Object Context(# SharedAttributes) UniqueMessages [0,0,0,3] 0 “[6,2,10,14,0]” “[6,2,14,10,0]” 1“[6,2,10,14,0]” “[6,2,14,10,0]” 2 “[6,2,10,1,0]” “[6,2,10,14,0]” “[6,2,10,5,0]” 3 “[6,2,10,5,0]” 6.6.3 Quantitative communication analysis Mappings between concepts and messages Third, we use information-theoretic scores and compare the context-unaware to the context-aware setting to quantify the results we obtained from our qualitative analysis. In the context-unaware setting, we observe high overall information scores (NMI scores ranging from 0.94 [0.9, 0.98] 9 for D(5,4) to 0.97 [0.96, 0.98] for D(3,8)). This suggests that concepts and messages tend to have one-toone mappings. While the mutual information between messages and concepts is also relatively high for context-aware agents, it is slightly lower than for context-unaware agents (NMI scores ranging from 0.86 [0.78, 0.92] for D(3,16) to 0.9 [0.88, 0.93] for D(3,8)). This could mean that context-aware trained agents adapt to the context, making strict one-to-one mappings impractical. Figure 6.4 shows for the context-unaware setting how the mutual information varies when it is calculated for all concept and context conditions for dataset D(4,4). 10 Here, we observe two patterns: On the one hand, the NMI increases with the number of fixed attributes. In other words, the more specific the concepts are, the more one-to-one mappings between concepts and messages emerge. On the other hand, the NMI scores stay relatively constant across different numbers of shared attributes. This suggests that context-unaware trained speaker agents adapt their choice of reference to a concept’s levels of abstraction, but not to the context (of which they are not aware). When looking at the NMI for the context-aware setting in Figure 6.5, we observe the opposite pattern: While changes in the concept level (i.e., the number of fixed attributes) are not reflected in changing NMI scores, we do observe increasing NMI scores with an increasing number of shared attributes. In other words, the finer the context, the more one-to-one mappings between concepts and messages can be found in the agents’ communication system. 9The intervals reported here are bootstrapped 95% Confidence Intervals. 10Plots for all datasets are available in Appendix D.
60 6 Case study 2: The role of context availability in emergent communication about concepts Figure 6.4: Context-unaware: Mean NMI scores across all datasets for different concept (# fixed attributes) and context conditions (# shared attributes). From top to bottom context becomes finer and from left to right concepts become more specific. Figure 6.5: Context-aware: Mean NMI scores across all datasets for different concept (# fixed attributes) and context conditions (# shared attributes). From top to bottom context becomes finer and from left to right concepts become more specific. Effect of the level of abstraction We will now look first at the effect of a concept’s level of abstraction and then at the effect of the context on the emerging language in more detail. The effect of a concept’s level of abstraction on the emerging language is visualized in Figures 6.6 and 6.7 which plot the entropy-based scores over different concept levels aggregated over all datasets and simulation runs for context-unaware and context-aware, respectively. In Figure 6.6, we observe that the NMI is largely constant for more specific concepts (three fixed attributes and more) and slightly drops toward more generic concepts with one or two fixed attributes. This effect is largely driven by a corresponding drop in the consistency score when it comes to more generic concepts, which suggests that more than one unique message is used to refer to the same generic concept. We can think of two reasons for this: One reason might be that the agents are overly specific when referring to the generic target concept, for example, they might use “red circle” or “blue circle” to refer to “circle”. Another reason is that the emerging language contains more synonymous words that refer to more generic concepts, for
6 Case study 2: The role of context availability in emergent communication about concepts 61 example the invented messages “1, 1, 2” and “2, 3, 4” both mean “circle”. Figure 6.6: Context-unaware: Mean entropy scores across all datasets for different concept levels indicated by the number of fixed attributes. From left to right concepts become more specific. Error bars indicate bootstrapped 95% confidence intervals. Figure 6.7 shows that we observe a drop in the consistency score when it comes to more generic concepts alsoforlanguagesdeveloped bycontext-awareagents.Additionally,wefind thatconsistency decreases again for more specific concepts (i.e., when the number of fixed attributes is larger than three). This can be explained by the availability of context in the context-aware setting: For more specific concepts with three or more attributes, there are more context conditions possible, i.e. 𝑛−1 context conditions. Thus, context-aware trained speakers adapt to use different messages to refer to the same concepts when they take context into account. The butterfly shape we observe in Figure 6.7, where effectiveness increases for specific and for generic concepts and consistency, on the other hand, decreases for specific and for generic concepts, can thus be explained by the two factors that the agents take into account when constructing messages, both concept specificity and context. Figure 6.7: Context-aware: Mean entropy scores across all datasets for different concept levels indicated by the number of fixed attributes. From left to right concepts become more specific. Error bars indicate bootstrapped 95% confidence intervals. We used Bayesian estimation to statistically analyze these observed differences between conditions across all five runs, following Kruschke (2013). We find no substantial difference in NMI scores between the context-unaware (M=0.92, CrI=[0.9, 0.93] 11 ) and the context-aware (M=0.89, CrI=[0.88, 0.91]) setting for generic concepts with an estimated difference in means of M=0.023 (CrI=[-0.003, 11Credible Intervals (CrIs) were computed on the posterior via the Highest Density Intervals.
62 6 Case study 2: The role of context availability in emergent communication about concepts 0.048], pd=95.9% 12 ,6%in ROPE 13 ). The difference for specific concepts between the context-unaware (M=0.94, CrI=[0.92, 0.97]) and the context-aware (M=0.87, CrI=[0.84, 0.91]) setting on the other hand is substantial with an estimated difference in means of M=0.07 (CrI=[0.026, 0.109], pd=99.4%, 0% in ROPE). While these effects are rather small, we do find reliable differences. These results are in line with our observations above, specifically that specific concepts can appear in a wider range of contexts (coarse to fine). Thus, context-aware agents use a wider range of messages to refer to the same specific concept than context-unaware agents because they can make use of the context. Effect of the context The effect of the context on the emerging language is especially evident when we compare Figure 6.8 and Figure 6.9 which plot the entropy-based scores over different context conditions for contextunaware and context-aware settings. In the context-unaware setting, the NMI stays at a constant level across different context conditions. We observe a small drop in consistency and an increase in effectiveness for fine contexts (i.e., for 3 or 4 shared attributes) in the datasets with at least 4 attributes. These results are in line with the hypothesis that context-unaware speakers communicate concepts on the most specific level in all contexts, including coarser contexts. This behavior can be referred to as overinformative from the listener’s perspective. For example, in a coarse context where no other circles are present, communicating a specific concept like “red circle” is considered overinformative. Figure 6.8: Context-unaware: Mean entropy scores across all datasets for different context conditions indicated by the number of shared attributes. From left to right context becomes finer. Error bars indicate bootstrapped 95% confidence intervals. When agents are trained context-aware, on the other hand, we observe that the information-theoretic scores differ more between context conditions (see Figure 6.9). Specifically, we observe a pattern where the coarser the context (i.e., the fewer shared attributes), the lower the NMI and the finer the context (i.e. the more shared attributes), the higher the NMI. When agents develop fewer one-to-one mappings between messages and concepts in the coarse context conditions, this might indicate 12 The probability of direction (pd) can be interpreted as the probability that a parameter’s posterior distribution is strictly positive or negative (Makowski, Ben-Shachar, Chen, & Lüdecke, 2019). 13 The Region Of Practical Equivalence with zero (ROPE) was calculated by using one-tenth of the standard deviation of the response variable around the null following recommendations by Kruschke, 2018: ROPE = [-0.004, 0.004].
6 Case study 2: The role of context availability in emergent communication about concepts 63 that they adapt more to the context which makes one-to-one mappings impractical. The reason for this might be that in coarse contexts, both more and less specific messages can be successful (e.g., “circle” can mean ‘red circle’, ‘blue circle’ etc.) because when less specific messages are used, the target concept can still be disambiguated by the context. In fine contexts, on the other hand, the messages need to contain more information on more specific levels of abstraction to be sufficiently discriminative in the context, which intuitively results in more one-to-one mappings (e.g., a more specific utterance like “red circle” is only used for the more specific concept ‘red circle’). Figure 6.9: Context-aware: Mean entropy scores across all datasets for different context conditions indicated by the number of shared attributes. From left to right context becomes finer. Error bars indicate bootstrapped 95% confidence intervals. In line with these observations, we find a substantial difference in NMI scores between the contextunaware (M=0.95, CrI=[0.94, 0.96]) and context-aware setting (M=0.89, CrI=[0.87, 0.9]) only for coarse contexts with a difference in means of M=0.064 (CrI=[0.046, 0.811], pd=100%, 0% in ROPE). For fine contexts, the difference in NMI scores between the context-unaware (M=0.95, CrI=[0.94, 0.97]) and the context-aware setting (M=0.95, CrI=[0.92, 0.97]) is not significant (M=0.008, CrI=[-0.026, 0.041], pd=70.1%, 20% in ROPE). Looking at effectiveness and consistency scores in the context-aware setting, we observe higher consistency and lower effectiveness scores for coarse contexts and higher effectiveness and lower consistency scores for fine contexts. This means that agents tend to consistently use the same messages to refer to the same concepts (i.e. no synonyms) in coarser contexts and that agents tend to effectively use messages that uniquely identify the target concept (i.e. non-polysemous expressions) in finer contexts. This makes sense because the finer the context gets, the more it is necessary to distinguish the target concepts from the distractors. 6.7 Discussion With our interactive agent-based model, we were able to generate three main insights about concept communication in various contexts and how this setup shapes an emerging language.
64 6 Case study 2: The role of context availability in emergent communication about concepts First, we show that artificial agents can learn to communicate successfully about concepts at different levels of abstraction and in different contexts in a concept-level reference game. Previous work has explicitly encoded concept information in the form of relevance vectors (Ohmer, Duda, & Bruni, 2022) or prototype embeddings (Mu & Goodman, 2021). For humans, however, abstracting the relevant concept, or level of reference, happens without such explicit information. Here, we show that agents can learn higher-level concepts from the object inputs alone, providing a more natural model for the emergence of abstraction. Second, we find that only context-aware agents learn to communicate efficiently by adapting their messages to the context conditions. While context-unaware agents use the same messages to refer to concepts in all context conditions, context-aware agents adapt their messages successfully to the context. Overinformative communication, in the sense that specific concepts are communicated also in coarse contexts where they contain more information than necessary for disambiguation, is reduced in the context-aware game scenario. This might indicate that context-aware agents communicate more efficiently (Piantadosi et al., 2012). It should be noted, though, that these agents do not share the same biases as humans. Future work should focus on the biases and pressures that shape the emerging language between artificial agents towards the kind of efficient overinformative communication we often observe in humans (e.g., Degen et al., 2020; Kreiss et al., 2017; Rubio-Fernandez, 2021; Tourtouri et al., 2019). Third, we conclude that the availability of context alone shapes the emerging language towards being more efficient (i.e. less overinformative) without additional pressures. The agents were not explicitly incentivized to use the context but they share the same architecture and training procedure with the context-unaware agents, the only difference being that they also receive distractor objects as input. Because we have not incentivized the context-aware agents to use context, they could follow the same strategy as context-unaware agents and be maximally specific all the time. Instead, we find that the agents develop a strategy that makes use of the context in which they communicate. Although the differences we observe between the context-aware and context-unaware settings are rather small, they are reliable and they do indicate that the mere presence of context already drives its use in communication. Future work can investigate whether pressures, such as increasing cognitive load for longer messages, would even intensify these differences. Our results are in line with previous work on how an emerging vocabulary depends on the contexts in which the targets are presented. Hawkins et al. (2018) found a similar pattern in an artificial language learning paradigm with human participants: The finer the context, the more one-to-one mappings are established in an emerging language, and the coarser the context, the more synonyms can be found. Further, they also found that an emerging language contains more words that refer to only one concept and fewer that refer to more than one concept when participants only encounter fine contexts. Our modeling results add to this evidence and highlight the role of context from a different angle. We treat neural network models as testbeds for hypotheses on human cognition. Here, we show that context in itself is a pressure that drives efficiency in an evolving language. Even though
6 Case study 2: The role of context availability in emergent communication about concepts 65 our neural network agents lack human cognitive biases, they develop more efficient protocols when they can (but do not have to!) access information about the context compared to when they cannot. This finding demonstrates that the presence of context alone may drive aspects of pragmatic communication without any additional pressures and cognitive prerequisites. We can take this as evidence for the role of external factors such as context for the emergence of an efficient communication system. In line with that, Piantadosi et al. (2012) argue that ambiguity, as we see it in the emerging communication system in the context-aware setting, makes a language efficient because it can usually be resolved by context. Our simulations provide evidence for this hypothesis. In conclusion, the here presented models and analyses contribute to our understanding of referential communication and the role of pragmatics in communicating concepts through a systematic manipulation of communicative needs. Our results show that the speaker’s access to the context shapes the emerging communication system, reproducing a pattern that was observed in humans (e.g., Hawkins et al., 2018; Winters et al., 2015,2018). These findings have implications both for linguistics research with the questions of how human language evolved and how we make use of language efficiently, as well as for emergent communication research with the question of how we can build artificial models that communicate in a human-like way. More generally, our work illustrates how language emergence simulations with neural network agents can be used to explore questions about human cognition. Acknowledgements We thank three anonymous reviewers for their helpful comments and feedback. The simulations were run on a high-performance computing cluster funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 456666331. Kristina Kobrock is supported by the DFG-funded Research Training Group “Computational Cognition” (DFG-GRK 2340). Author Contributions: Kristina Kobrock: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Writing - Original Draft, Visualization. Xenia Ohmer: Conceptualization, Methodology, Software, Writing - Review & Editing, Visualization. Elia Bruni: Conceptualization, Methodology, Writing - Review & Editing, Supervision. Nicole Gotzner: Conceptualization, Methodology, Writing - Review & Editing, Supervision, Project Administration.
66 6 Case study 2: The role of context availability in emergent communication about concepts Appendix A Accuracy scores across all datasets Table 6.4: Accuracy means for agents trained in the context-unaware and context-aware setting averaged over five runs with standard deviations. Datasets Condition Accuracy training validation test D(3,4) context-unaware 0.995 (0.002) 0.99 (0.003) 0.84 (0.036) context-aware 0.993 (0.003) 0.983 (0.004) 0.784 (0.035) D(3,8) context-unaware 0.993 (0.003) 0.989 (0.003) 0.778 (0.068) context-aware 0.984 (0.006) 0.977 (0.006) 0.686 (0.061) D(3,16) context-unaware 0.981 (0.007) 0.979 (0.008) 0.896 (0.005) context-aware 0.969 (0.005) 0.968 (0.006) 0.874 (0.007) D(4,4) context-unaware 0.992 (0.002) 0.989 (0.002) 0.922 (0.028) context-aware 0.995 (0.003) 0.993 (0.005) 0.942 (0.048) D(4,8) context-unaware 0.961 (0.011) 0.961 (0.011) 0.943 (0.012) context-aware 0.984 (0.004) 0.982 (0.006) 0.976 (0.007) D(5,4) context-unaware 0.98 (0.011) 0.979 (0.012) 0.964 (0.014) context-aware 0.985 (0.007) 0.984 (0.008) 0.979 (0.01)
6 Case study 2: The role of context availability in emergent communication about concepts 67 B Errors across all datasets B.1 Errors per game round These plots show the errors on the validation dataset across all datasets for different concept (# fixed attributes) and context conditions (# shared attributes). Game rounds in which at least one object was incorrectly classified count as errors and are normalized with the number of occurrences of the specific condition in the dataset. This means that a value of 1.0 indicates that listeners incorrectly classified at least one object in each game round in this condition. Figure 6.10: Context-unaware: Most errors occur on the diagonal from top left to bottom right, i.e. in the finest possible context conditions. As can be seen in the figures, errors occur mostly in fine context conditions, i.e. where the maximally possible number of attributes is shared between targets and distractors. Some of these errors are false positives, i.e. distractors are incorrectly classified as targets, and some of these errors are false negatives, i.e. targets are incorrectly classified as distractors. We find that false negative errors occur mainly with more generic concepts and fine contexts. This is probably due to the target concepts being very heterogenous and thus, harder to discriminate against distractors. False positive errors, on the other hand, occur in the finest contexts when the concept is very specific. This can be explained by the distractors being very similar to the targets in these conditions. In other words, false positive errors might indicate that the learned target concept is a bit too wide, and false negative errors might indicate that the learned target concept is a bit too narrow.
74 6 Case study 2: The role of context availability in emergent communication about concepts 6.8 Summary of the key contributions In this case study, we provided evidence for the role of pragmatics in the evolution of a communication system. To study this evolution, we employed an agent-based emergent communication modeling paradigm aimed at achieving maximal control and systematicity. We introduced carefully constructed symbolic datasets that enabled us to manipulate both conceptual and contextual informativity. Specifically, we investigated how agents communicate target concepts with varying specificity ranging from specific to generic, in contexts with varying granularity ranging from fine to coarse. Going beyond previous work (Mu & Goodman, 2021; Ohmer, Duda, & Bruni, 2022), we demonstrated that agents can ascertain target concepts based solely on object-based inputs, without any additional explicit information. Our key finding from comparing speakers with access to the context to speakers without access to the context was that the mere presence of context drives its use in communication. We further found that communication strategies differ depending on whether speakers had access to the context during training or not. Context-aware agents tend to vary their messages more in response to context. They use a broader set of messages in coarse contexts, whereas in fine contexts, the most successful strategy is to communicate the most specific concept. This suggests that in coarse contexts, speakers are more likely to come up with multiple names for the same concept. Context-unaware speakers, in return, need to communicate all attributes that are relevant to the target concept disregarding the context. While this can be considered an effective semantic solution, it may also be overinformative from the listener’s perspective, who interprets the messages within context. Although these context-unaware languages are sensitive to the levels of the conceptual hierarchy ranging from specific to generic, they lack context-dependent differences in reference. In other words, here, agents do not come up with multiple names for the same concept, but rather use different names for each concept. We conclude that the pragmatic factor of context plays a crucial role in shaping the emergent communication system, contingent on its availability to the speaker. 6.9 Implications for the broader research question We introduced a computational model to investigate the emergence of language during interactive communication about concepts. This model will also be used in the following two case studies. One goal of this study was to introduce the model and dataset, demonstrating how they help us investigate the role of pragmatic factors on language evolution. Our results indicate that artificial agents can learn to communicate about the hierarchically organized concepts in our symbolic datasets. Notably, this is achieved without any explicit encoding of information about the conceptual hierarchy. We thus move beyond previous research where such explicit information was provided in the form of relevance or prototype vectors (Mu & Goodman, 2021; Ohmer, Duda, & Bruni, 2022). We can conclude from these results that the concept-level reference game is an effective task for highlighting the conceptual structure of the input data. This is consistent with the findings from case study 1, in which we found that the concept-level reference game enables participants to traverse the
6 Case study 2: The role of context availability in emergent communication about concepts 75 conceptual hierarchy and flexibly select appropriate referring expressions from different levels of this hierarchy. The second goal of this study was to investigate the role of pragmatics in the evolution of a communication system. We showed that the availability of context influences how concepts at different levels of abstraction are referred to in a language. Previous experimental work using an artificial language learning paradigm has also provided evidence for the role of context in shaping communication systems (Hawkins et al., 2018; Winters et al., 2015,2018). However, our research highlights that even without explicit incentives to utilize context, its availability drives its use in communication, thereby shaping the emerging language. Furthermore, we provided initial evidence that languages which emerged when context was available are more efficient than languages which emerged when context was not available to the speakers. We will follow up on these findings in case study 3. Additionally, we discovered that the conceptual hierarchy and the context granularity are reflected differently in properties of the language, depending on the context availability. In the context-unaware condition, languages more closely reflect the conceptual structure. We can relate this to the semantics of a language where the lexicon consists of mappings between concepts and words. Adding pragmatics in the form of context in the context-aware condition, leads to languages reflecting the context granularity. This corresponds to a language where referring expressions can be flexibly and pragmatically adapted to different contexts. Importantly, all observations are emergent features of the model, arising from the structure of the input data and the communicative needs during interactions, without any explicit incentives for the agents to encode concept specificity or context granularity. Thus, we conclude that the distinctive linguistic properties we observed emerged from the environment and communicative needs.
7 Case study 3: The role of pragmatic mechanisms in language use and the emergence of linguistic and conceptual structure This chapter presents case study 3. The chapter starts with a high-level introduction followed by the content of the publication: Kobrock, K., Ohmer, X., Bruni, E., & Gotzner, N. (2025b). The role of pragmatic mechanisms in referential communication and categorization: An emergent communication model. [under review, PsyArXiv]. https://doi.org/10.31234/osf.io/kbfua_v1 The chapter ends with a brief summary of the main contributions of the publication and a discussion of its implications for the broader research question of this dissertation. 7.1 High-level introduction Case study 2 provided initial evidence for the role of language and pragmatics in conceptual abstraction and, specifically, in shaping an evolving language that is used for communicating concepts at different levels of abstraction. The goal of case study 3 is to investigate this role further. We explore the possibility of a joint evolution of language and category systems through two experiments. In Experiment 1, we aim to investigate the influence of a shared context on language emergence. In Experiment 2, we investigate the role of different pragmatic mechanisms on language use in concrete reference situations. For this purpose, we change four main aspects with respect to the model used in the previous case study: First, we introduce a shared context between speakers and listeners that is more informative than the context modeled in case study 2. In the previous model, only the number of attributes shared between target and distractor objects was considered when constructing the context. Here, speakers and listeners see the target concepts in the same context, ensuring that the attributes that are shared with the target concept are the same for both interlocutors. For instance, the speaker needs to distinguish small blue triangles from small red triangles. In this example, the target concept has three attributes that need to be communicated, small, blue, and triangle. It is presented in a fine context, where two of the attributes, small and triangle, are shared between targets and distractors. The discriminative attribute is thus the color. Likewise, the listener needs to discriminate small blue triangles from small triangles with a different color, for example small green triangles. This alignment ensures that the same attributes are relevant for communication. We hypothesize that introducing this shared context will make the context more useful for communication in comparison to case study 2, and this will, in turn, impact the emerging language. Second, to enhance the emergence of efficient category and linguistic systems, we add a length cost pressure to the training objective. This pressure penalizes longer messages by adding a cost to the training loss. Consequently, speaker-listener pairs receive a higher reward for interactions in which speakers use a shorter message than for those in which speakers
78 7 Case study 3: The role of pragmatic mechanisms in language use and emergence use a longer message, provided the communicative success, that is, the listener’s correct selection of the target objects, remains the same. We hypothesize that the length cost pressure will make the emerging language more efficient. Agents might be more inclined to communicate only the necessary attributes of a concept when communication is more costly. Third, in Experiment 2, we evaluate not only the mappings between messages and concepts that emerged as a function of training, but we evaluate language use after training. This is done by testing the agents on a test dataset consisting of held-out concepts. Investigating how agents use the languages they developed during repeated interactions with their interlocutors provides a new window of analysis for our research question. By integrating language use and language emergence, we can account for pragmatic factors on both evolutionary and situation-based time scales. Specifically, investigating language use in concrete reference situations will be of central importance for drawing inferences from the linguistic to the category system and will help us analyze how concepts at different levels of abstraction are communicated in reference-based situations. Finally, we add two new experimental conditions to manipulate not only speakers’ access to context but also their ability to consider the expected utility of their utterance by reasoning about the listener’s likely interpretation of each potential utterance. This is modeled via the RSA framework introduced in Section 3.2.RSA speakers aim to maximize the utility of their utterances, which is influenced by (1) the likelihood of the listener selecting the correct target objects, and (2) the cost associated with an utterance in terms of its length, where shorter messages yield higher utility. We expect agents in Experiment 1 to develop more efficient languages in the context-aware compared to the context-unaware condition consistent with case study 2. Regarding the comparison of language use in Experiment 2, we expect that both context availability and the availability of RSA-like reasoning about the utterances expected utility, will lead to speakers making more effective and efficient use of their language.
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 79 7.2 Abstract We model pragmatic mechanisms of referential communication and categorization in a multi-agent framework of emergent communication. Pragmatic theories and experimental work predict that speakersconsider the contextin their choiceof referringexpressions. Inaddition tothis context-based reasoning, utility-based pragmatic reasoning about the listener’s likely interpretation of an utterance influences the speaker’s production choices. We aim to investigate these two factors and their role in referring expression generation and categorization in a computational model of language emergence and language use. We model communication in interaction and consider an efficiency tradeoff between speaker and listener utilities. Our results show that an emerging language becomes more effective, ambiguous, and efficient when speakers and listeners communicate in a shared context. This is achieved by an efficient tradeoff between production and comprehension where languages can afford to be simpler in production when they are sufficiently informative in context, placing more burden on the listener’s side. We further demonstrate that incorporating utility-based pragmatics, as modeled with the Rational Speech Acts framework, improves the linguistic efficiency of language use only in languages that emerged with a shared context between interlocutors, but not in languages that emerged without such contextual information during training. We conclude that context-based pragmatics plays a role in referential communication and categorization by shaping an emerging language. Efficient reference in a communicative situation can benefit especially from utility-based pragmatics if the language that is being used has emerged in context. This might suggest that utility-based pragmatics hinges on mechanisms that naturally emerge when context is available during the evolution of a language. In summary, we show that human-like language and category systems emerge as an optimized tradeoff between speaker and listener needs in interaction, i.e. under efficiency considerations of simplicity and informativeness. Author summary This study investigates pragmatic mechanisms of referential communication and categorization in a multi-agent framework. We study how speakers and listeners interact to create an efficient language, balancing the need for clear communication with the cost of producing messages. Our results show that an emerging language becomes more effective, ambiguous, and efficient when speakers and listeners communicate in a shared context. This mechanism of context-based pragmatics plays a crucial role in shaping the language, allowing speakers to be simpler in production when listeners can infer the correct meaning in context. Additionally, we discovered that utility-based pragmatics modeled with the Rational Speech Acts framework impacts language efficiency, particularly when the language has emerged in context. Our findings suggest that utility-based pragmatics is closely tied to the mechanisms that naturally emerge when context is available during language evolution. The resulting language and category systems resemble those of human languages. This research contributes to our understanding of how efficient communication systems emerge in social interactions, and based on communicative need.
80 7 Case study 3: The role of pragmatic mechanisms in language use and emergence 7.3 Introduction The languages we speak influence which categories we know and talk about. The idea that language shapes cognition in this way has become popular as the (weak) Sapir-Whorf hypothesis, or linguistic relativity hypothesis (Sapir, 1912; Whorf, 1956). One instance of the claim that language shapes cognition is that conversation shapes category structure (Barr & Kronmüller, 2006). A famous example is the popularized claim from Boas (1911) and Sapir (1912) that Inuit 1 have more words for subcategories of snow than languages spoken in warmer climates. This claim has been recently revisited in a large-scale analysis of linguistic and meteorological data (Regier et al., 2016). The authors found that languages that use the same linguistic form for both snow and ice tend to be spoken in warmer countries. They relate this finding to the lower communicative need to talk about snow and ice in these regions (Regier et al., 2016). This is in line with the efficient communication hypothesis which predicts that human languages are optimized for efficient communication, a claim that has gained traction in many subfields of linguistics such as language evolution (Kanwal et al., 2017; Winters et al., 2018), pragmatics (Kreiss et al., 2017; Peloquin et al., 2020; Rubio-Fernandez, 2021; Rubio-Fernandez et al., 2021) and more general proposals (Gibson et al., 2019; Gualdoni & Boleda, 2024; Levshina, 2018; Piantadosi et al., 2012; Regier et al., 2016; Zipf, 1949). Following the efficient communication hypothesis, the main idea for how language shapes category structure is that languages have words for things that are relevant, i.e. where there is a communicative need (Kemp & Regier, 2012; Regier et al., 2016). This can be exemplified by the following relationship (Regier et al., 2016): Environment →Communicative Need →Category System (7.1) If there is a communicative need to talk about subcategories in an environment, then the category system will be more fine-grained than if there is no such need (Regier et al., 2016). More fine-grained categories are more informative than broader categories, but they are also more complex and require more storage space in memory. It has been proposed that, due to cognitive constraints, both languages and category systems tend to find an optimal tradeoff between simplicity and informativeness (Kemp & Regier, 2012; Piantadosi et al., 2012; Regier et al., 2016; Rosch, 1978; Rosch et al., 1976; Winters et al., 2018; Zipf, 1949). How this tradeoff is optimized can be better understood when considering interactive communication. Let’s assume that a lexical system evolves in a simple conversational setup in which a speaker describes an object to a listener. If this system was optimized for the speaker, then in an extreme case, the system would only include one word that can be used to describe all objects. This optimizes simplicity and keeps memory storage for the speaker to a minimum. A system that is optimized for the listener, on the other hand, includes one label for each object. This optimizes informativity of the lexicon and it makes it easy for the listener to identify the correct target when hearing a word. This principle was proposed as the Principle of Least Effort by Zipf (1949) and the 1 The early works and debate thereafter have frequently used the term ‘Eskimo’ which is largely dispreferred by the native populations of Alaska due its colonial heritage (Kaplan, 2025).
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 81 idea has been also taken up later (Levinson, 2000; Piantadosi et al., 2012). Relatedly, Rosch et al. (1976) and Rosch (1978) have proposed a similar principle for category structure: The Principle of Cognitive Economy states that a category system should provide maximum information with the least cognitive effort. An optimal category system would thus optimize a tradeoff between having many categories with very fine-grained distinctions between them and having fewer categories with coarser distinctions between them. A category system with many categories is optimized for informativity: Many properties of one category can be predicted by knowing only one property, e.g. the label. A category system with few categories is optimized for cognitive economy because fewer categories need to be learned and stored in memory (Rosch, 1978; Rosch et al., 1976). This tradeoff can also be thought of as a tradeoff between speaker and listener needs in a communicative situation, specifically in a referential communicative situation. Barr and Kronmüller have argued that referential communication is especially critical for categorization, specifically category learning and category use (Barr & Kronmüller, 2006). Following this, we use referential communication to study categorization in a communicative setting where speakers and listeners interact. This allows us to study communicative and categorization efficiency as a tradeoff between speaker and listener needs. The question that we tackle in this paper is the following: Can the efficient communication hypothesis predict category structure? Evidence for this has been found in the context of kinship (Kemp & Regier, 2012) and color (Fider & Komarova, 2019; Ohmer, Marino, et al., 2022; Zaslavsky et al., 2018, 2022) systems. However, this hypothesis has not yet been tested for more general categories or even for the most ubiquitous category system that we use: How we categorize the things around us in hierarchical structures that are also called levels of abstraction (Rosch et al., 1976), e.g. dalmatians, dogs and animals. To study categorization, we look at the referential communication of categories at different levels of abstraction. We use a modeling approach with controlled symbolic datasets that include hierarchical categories. One central puzzle for theories of efficient communication is how to account for ambiguity and redundancy observed in human language (Degen et al., 2020; Piantadosi et al., 2012; Wasow, 2015). Ambiguity in human languages has been related to pragmatic considerations (Achimova et al., 2022; Ferreira, 2008; Piantadosi et al., 2012). Pragmatic theories assume that language is interpreted in context and in a referential situation where speakers use an underinformative or ambiguous expression, context usually helps to disambiguate the intended meaning (Piantadosi et al., 2012; Wasow, 2015). Context-based pragmatics can thus reconcile ambiguity with the efficient communication hypothesis. But, humans have also been shown to frequently overspecify referring expressions with atypical or salient properties such as color even in situations where these properties are not needed to unambiguously identify a target referent in a given context. On first sight, this is at odds with the efficient communication hypothesis and Gricean maxims of conversation, especially the quantity maxim which states that speakers should provide as much information as needed and not more (Grice, 1975). But recent proposals have shown that the redundant expressions provide information that helps listeners to identify the target referent, e.g. by mentioning atypical features or providing salient information in complex scenes that facilitate visual search (Degen
82 7 Case study 3: The role of pragmatic mechanisms in language use and emergence et al., 2020; Kreiss et al., 2017; Rubio-Fernandez, 2021; Rubio-Fernández, 2016; Tourtouri et al., 2019). This can be explained by the idea that speakers take into account the listener’s needs and want to guide the listener toward identifying the correct referent as quickly as possible (Degen et al., 2020; Tourtouri et al., 2019). Reasoning about the listener’s likely interpretation is a classic case of utility-based pragmatics and it can be straightforwardly modeled with the Rational-Speech-Act framework (Degen et al., 2020; Frank & Goodman, 2012; Franke & Degen, 2016). Utility-based pragmatics can thus explain how redundancy can be considered efficient and reconcile redundancy with the efficient communication hypothesis. The main objective of this work is to find out how categorization influences an emerging language and what role language and pragmatics play in categorization. The goals of our study are thus twofold: First, we investigate the influence of our manipulations on the (use of the) emerging language. We ask whether efficient communication, formalized as a tradeoff between speaker and listener needs, emerges from interaction and pragmatic considerations, when communicating concepts at different levels of abstraction. We use an emergent communication paradigm to model the emergence of categories and category labels in an interactive communicative framework that takes into account speaker and listener needs. Second, we draw inferences from the emerged linguistic system to the category system and analyze how concepts at different levels of abstraction are communicated in the language. We use the trained emergent communication models and test them in specific communicative situations and under consideration of different pragmatic mechanisms, such as context-based and utility-based pragmatics. 7.3.1 Modeling approach The overarching goal of this research is to model context-based and utility-based pragmatic mechanisms in a communicative interaction that involves categorization, and to find out under which circumstances they become especially useful and make the (use of the) emerging language particularly efficient. Our research makes use of an emergent communication model (Kobrock, Ohmer, et al., 2024; Lazaridou et al., 2017,2018; Ohmer, Duda, & Bruni, 2022). The architecture of the model is visualized in Fig7.1 and implements four communicative principles that are outlined below. First, the model is interactive, i.e. it takes into account and models the fact that communication happens in interaction between a speaker and a listener agent. This means that our model should have at least a basic one-way interaction between a speaker who sends a message and a listener who receives a message. The model that we use utilizes such a basic interaction which is also utilized in standard linguistic theoretical and experimental frameworks (Frank & Goodman, 2012; Hawkins et al., 2018). Second, the model should make as little prior assumptions about communication as possible. Therefore we do not use a model of language use on an existing language, for example English, but rather let a proto-language emerge in interaction between artificial agents. The goal is to observe
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 83 Speaker neural network (GRU) Listener neural network (GRU) message [4,11,7,0] prediction and training 1 2 3 4 Figure 7.1: Model architecture. what a language might look like given the constraints we implement. If in our agent-based models we observe that a language emerges that shares certain characteristics with human language, we can conclude that these characteristics can emerge even in simple neural-network-based models that implement some central characteristics of communication, such as a reference-based interaction, but do not implement other uniquely human constraints, such as language universals or a particular grammar (Chaabouni et al., 2021). Third, the model takes into account that communication should be grounded. For grounded communication, it is crucial that communication is situated in a world - this can be a toy world to keep the complexity of the model small. This is done by simplifying communication to referring. The agents need to solve a reference game where the speaker has to describe a target and the listener has to select the target from a set of distractors. The emerging language is a meaningful mapping between messages and objects of the world that are being referred to. This simple task and setting have an additional advantage: We have full control over the environment, i.e. we can construct the properties of the objects and thereby control for every parameter of the model: The objects that should be referred to and the context in which the reference takes place. In other words, when controlling for the targets and the context, we can control for how much information needs to be communicated to achieve successful communication. The amount of information that needs to be communicated relates closely to efficiency. Fourth, the model should allow us to investigate categories at different levels of abstraction. To achievethisgoal,weimplementaconcept-levelreferencegamethat the agents need tosolve(Kobrock, Ohmer, et al., 2024; Kobrock, Uhlemann, & Gotzner, 2024; Mu & Goodman, 2021). The concept-level reference game is similar to a simple reference game. The only difference is that instead of a single target object, a target concept needs to be communicated. This target concept is constructed by combining several target objects. For example, it has been shown in experimental work that when speakers observe two targets, a parrot and a dog, they use the superordinate referring expression “animal” to refer to both (Kobrock, Uhlemann, & Gotzner, 2024). The concept-level reference game ensures that the agents in our model communicate about concepts at different levels of abstraction
90 7 Case study 3: The role of pragmatic mechanisms in language use and emergence (a) Context-unaware (b) Context-aware with shared context Figure 7.4: Message lengths per concept hierarchy level. Concept specificity, or the amount of information that needs to be communicated, increases with the number of fixed attributes that is shared among target objects from more generic concepts (with fewer fixed attributes) on the left, to more specific concepts (with more fixed attributes) on the right. Bayesian hierarchical models predicting the respective score (NMI, effectiveness, consistency) by the condition. We found a main effect for condition in all three models, i.e. predicting NMI (M=0.14, CrI=[0.13, 0.16], pd=100%, ROPE=[-0.02, 0.02], 0% in ROPE), predicting effectiveness (M=0.22, CrI=[0.20, 0.25], pd=100%, ROPE=[-0.02, 0.02], 0% in ROPE), and predicting consistency (M=0.04, CrI=[0.03, 0.05], pd=100%, ROPE=[-0.01, 0.01], 0% in ROPE). This means that the amount of one-to-one mappings between messages and concepts was lower, and that the emerging language contained more synonyms and polysemous messages, in the context-aware condition compared to the context-unaware baseline. Table 7.3: Mean entropy-based scores, i.e. NMI, effectiveness and consistency. Context-unaware Context-aware shared context Datasets NMI effectiveness consistency NMI effectiveness consistency (3,4) 0.81 ±0.01 0.90 ±0.05 0.74 ±0.01 0.59 ±0.03 0.52 ±0.04 0.68 ±0.01 (3,8) 0.87 ±0.00 0.99 ±0.00 0.77 ±0.00 0.58 ±0.01 0.47 ±0.02 0.75 ±0.00 (3,16) 0.86 ±0.04 0.83 ±0.10 0.90 ±0.07 0.59 ±0.09 0.48 ±0.11 0.77 ±0.01 (4,4) 0.85 ±0.00 0.97 ±0.02 0.77 ±0.01 0.54 ±0.03 0.44 ±0.04 0.69 ±0.01 (4,8) 0.86 ±0.02 0.89 ±0.05 0.84 ±0.05 0.55 ±0.06 0.45 ±0.07 0.73 ±0.01 (5,4) 0.83 ±0.08 0.88 ±0.16 0.79 ±0.01 0.51 ±0.02 0.41 ±0.02 0.69 ±0.01 Mean entropy scores and standard deviations over five runs are calculated for each dataset and both conditions. How is it that the context-aware agents had such an ambiguous language but still achieved very high performance? We looked at entropy-based scores calculated for each level of the conceptual hierarchy to address this question. Fig 7.5A-B shows the information-theoretic scores comparing the context-unaware to the context-aware condition. The effectiveness score is closely related to communicative success, i.e. how likely the listener is to select the correct target objects. If effectiveness is high, this means that meaning uncertainty is low. Effectiveness is generally higher in the context-unaware than in the context-aware condition as shown in Table 7.3 and in Fig 7.5A-B. Interestingly, in both context-aware and context-unaware conditions, generic concepts with only
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 91 one fixed attribute were referred to almost with maximal effectiveness. Thus, there was no meaning uncertainty for messages that were used to refer to generic concepts, i.e. they effectively singled out the targets. The effectiveness score had the tendency to drop the more specific a concept was which means that there were more polysemous messages and meaning uncertainty the more specific a concept. While in the context-unaware condition the effectiveness score dropped gradually and had a linear trend, in the context-aware condition, the effectiveness score dropped exponentially with each level of the conceptual hierarchy. The consistency score can be viewed as a measure for of the speaker’s tendency to select the same message for a given target concept in repeated iterations. If the consistency score is high, this means that there is low signal uncertainty. In other words, the speaker’s choice of message is straightforward and easy. In the context-unaware condition, we observed higher consistency, i.e. lower signal uncertainty, the more specific the concepts. In the context-aware condition, we observed lower consistency, i.e. higher signal uncertainty, the more specific the concepts. The NMI curves of both conditions reflected the described trends: In the context-unaware condition, the NMI score was more closely oriented to the consistency score, i.e. there were more one-to-one mappings the more specific a concept. In the context-aware condition, NMI followed the effectiveness curve and decreased with increasing number of fixed attributes, suggesting that there were fewer one-to-one mappings between messages and concepts with specific than with generic concepts. The above described trends were supported by three hierarchical Bayesian models predicting the respective scores by the interaction of condition and concept hierarchy level, i.e. fixed attributes. We found a main effect for condition and concept hierarchy level, as well as an interaction effect for these two predictors in models predicting the NMI, effectiveness and consistency (see Table 7.9 in S2 Appendix: Statistical models). (a) Context-unaware (b) Context-aware with shared context Figure 7.5: NMI, consistency and effectiveness scores for each level of the conceptual hierarchy. Discussion In summary, we observed that a more ambiguous and efficient language emerged in the contextaware condition than in in the context-unaware condition which is in line with our predictions. This finding was supported by measures of communicative success (accuracy), of efficiency (message length), and of ambiguity (entropy-based scores). We found that communicative success on the validation dataset was substantially higher in the context-aware than in the context-unaware
92 7 Case study 3: The role of pragmatic mechanisms in language use and emergence condition. Messages were substantially shorter in the context-aware than in the context-unaware condition. Together with the high performance, this means that the languages that emerged in the context-aware condition are likely simpler which makes them well interpretable and well suited for generalization purposes. When interpreting the entropy-based scores we obtained for each language, it was evident that the languages that emerged in the context-aware condition were more ambiguous. We found fewer one-to-one mappings between messages and concepts shown by the substantially lower NMI score in the context-aware compared to the context-unaware condition. The substantially lower effectiveness scores in the context-aware condition mean that there is high meaning uncertainty and that listeners have to choose the most likely interpretation from more than one candidate interpretation. Together with the high accuracies, this suggests that listeners successfully make sense of the polysemous messages they receive. In the context-unaware condition, effectiveness scores were quite high, suggesting that the messages that the listeners receive were not ambiguous and it is thus straightforward for the listeners to interpret the messages and identify the correct targets. The consistency scores suggested that there is a certain amount of signal uncertainty, or synonymous messages, in the emerging languages of both conditions with little, but substantially, more synonymous messages in the context-aware than in the context-unaware condition. The effectiveness and consistency scores can also provide insight into how the languages emerged as a tradeoff between speaker and listener needs. The efficient communication hypothesis predicts that an efficient language for the speaker will have as few distinct messages as possible that can be used in multiple contexts. This translates into a high consistency score, i.e. low meaning uncertainty. An efficient language for the listener, on the other hand, will have many distinct messages because that makes interpretation easier. This translates into a high effectiveness score, i.e. low meaning uncertainty. In the context-unaware condition, we observed higher effectiveness than consistency. This means that the emerging languages in this condition were more optimal for listeners than for speakers. The languages have messages that can be interpreted unambiguously which makes the listener’s task especially easy. In other words, the tradeoff between production and comprehension is balanced by putting burden on the speaker who has to select a message from a set of multiple possible utterances and send quite long messages that make target identification easy for the listeners. In the context-aware condition, we observed the opposite pattern, namely that consistency is higher than effectiveness. This puts more burden on the listener’s side who has to interpret the meaning of polysemous messages in context. The speaker, on the other hand, has a lower burden than the listener. This is due to the lexicon having evolved towards a simpler lexicon with polysemous signals. This ambiguity makes the resulting language more efficient. This is only possible in the context-aware condition because the context is informative and can help the listener to resolve the ambiguity of the polysemous messages. In the referential communication task we employ, agents have to communicate about concepts ranging from specific to generic presented in a context of varying granularity. Looking at how concepts of varying specificity are encoded in the emerging languages provides insights into the category structure learned by the agents. We have conducted a concept-level analysis of the entropy-based scores to assess this. In both conditions, we observed that generic concepts with only one fixed attributes are referred to with almost maximal effectiveness, i.e. very low meaning
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 93 uncertainty. Moving towards more specific concepts, the effectiveness drops gradually in the context-unaware condition and exponentially in the context-aware condition. The exponential drop in the context-aware condition reflects the fact that the number of possible contexts, and thus the possibility to exploit these contexts for efficient communication, increases exponentially, the more attributes are fixed in a target concept. This trend can be also observed in natural languages where the same specific concept can be referred to with different labels, e.g. “animal”, “dog”, or “dalmatian”, depending on the context. The most generic concepts, on the other hand, e.g. animal or thing, tend to have only one label. We conclude that the availability of a shared context not only makes the emerging language more efficient, but also shapes the category system and how concepts at different levels of abstraction are referred to. 7.4.2 Experiment 2: Utility-based pragmatic reasoning The goal of the second experiment was to investigate the influence of utility-based reasoning on the situational reference to concepts at different levels of abstraction in different contexts. We looked at situational reference after a language has emerged in the context-unaware and context-aware conditions from the first experiment. This means that we investigated the interactions on a test dataset with concepts that have not been included in the train and validation datasets. First, we compared accuracies collected on the test dataset for context-unaware agents, contextunaware trained agents with RSA, context-aware agents, and context-aware trained agents with RSA. Table 7.4 shows that test accuracies were higher in the context-aware than in the contextunaware condition. This means that agents generalized better in the context-aware condition. Regarding utility-based pragmatic reasoning, accuracies in the context-unaware condition did not improve with RSA but RSA seemed to worsen performance of context-unaware agents. In the context-aware condition, however, utility-based pragmatics lead to similar performance. These effects were supported by a Bayesian hierarchical model predicting test accuracies by the interaction of context-based (context-unaware vs. context-aware) and utility-based (without RSA vs. with RSA) pragmatics. We found a main effect for context-based pragmatics (M=-0.08, CrI=[-0.09, -0.07], pd=100%, ROPE=[-0.01, 0.01], 0% in ROPE), a main effect for utility-based pragmatics (M=0.03, CrI=[0.02, 0.04], pd=100%, ROPE=[-0.01, 0.01], 0% in ROPE) and an interaction effect (M=0.03, CrI=[0.02, 0.04], pd=100%, ROPE=[-0.01, 0.01], 0% in ROPE). Table 7.4: Accuracies on the test dataset. Dataset Context-unaware + RSA Context-aware + RSA (3,4) 0.79 ±0.02 0.87 ±0.04 0.93 ±0.02 0.96 ±0.02 (3,8) 0.87 ±0.03 0.82 ±0.03 0.97 ±0.01 0.99 ±0.01 (3,16) 0.95 ±0.04 0.72 ±0.03 0.96 ±0.07 0.95 ±0.09 (4,4) 0.88 ±0.02 0.75 ±0.04 0.98 ±0.01 0.98 ±0.01 (4,8) 0.89 ±0.07 0.66 ±0.02 0.97 ±0.02 0.96 ±0.04 (5,4) 0.88 ±0.09 0.67 ±0.07 0.99 ±0.01 0.98 ±0.03 Mean accuracies and standard deviations over five runs are calculated for each dataset.
94 7 Case study 3: The role of pragmatic mechanisms in language use and emergence Next, we looked at the efficiency of the messages in relation to the amount of information that needed to be communicated, i.e. the number of fixed attributes in a concept (Fig 7.6A-D). In the context-unaware condition (Fig 7.6A), the mean length of messages used to refer to generic concepts was 10.5 and the mean length of messages used to refer to specific concepts was 17.2. When adding RSA (Fig 7.6B), messages in the context-unaware condition became much shorter, i.e. they used messages of 5.5 symbols on average when referring to generic concepts and messages of 10 symbols on average when referring to specific concepts. In the context-aware condition (Fig 7.6C), message lengths were very short in general. In reference to generic concepts, context-aware agents used messages with 3.6 symbols on average and messages with 3.9 symbols on average in reference to specific concepts. Adding RSA reduced the mean lengths of messages used by context-aware trained agents by about 1 symbol in reference to both specific and generic concepts (Fig 7.6D). Overall, RSA improved efficiency and lead to shorter messages being chosen by the speaker agents. These observations were in line with the predictions of a Bayesian hierarchical model predicting message length by the factors context-based pragmatics (context-unaware vs. context-aware), utility-based pragmatics (without RSA, vs. with RSA) and conceptual hierarchy level (number of fixed attributes). We found a main effect for context-based pragmatics (M=1.88, CrI=[1.28, 2.46], pd=100%, ROPE=[-0.52, 0.52], 0% in ROPE), a likely existing main effect with undecided significance for utility-based pragmatics (M=0.70, CrI=[0.09, 1.30], pd=98.78%, ROPE=[-0.52, 0.52], 27.82% in ROPE), an interaction effect between context-based pragmatics and conceptual hierarchy (M=0.87, CrI=[0.58, 1.18], pd=100%, ROPE=[-0.52, 0.52], 0% in ROPE) and a possibly existing three-way interaction between context-based pragmatics, utility-based pragmatics and conceptual hierarchy with undecided significance (M=0.28, CrI=[-0.02, 0.59], pd=97.02%, ROPE=[-0.52, 0.52], 97.05% in ROPE). Next, we looked at the lexical efficiency of the concept-message mappings. We calculated lexicon sizes and informativeness on the test interactions, i.e. the messages that have been produced when agents have been presented with concepts in the test dataset. Table 7.5 shows the number of concepts in the test datasets and the number of unique messages produced in the four different conditions. In a hierarchical Bayesian model predicting lexicon size by context-based and utility-based pragmatics and their interaction, we found no main effect for context-based pragmatics (M=-9.18, CrI=[-26.92, 8.37], pd=84.67%, ROPE=[-26.42, 26.42], 99.68% in ROPE), i.e. no difference between contextunaware and context-aware training with regards to the lexicon sizes. However, we observed a main effect of utility-based pragmatics (M=89.59, CrI=[71.74, 107.88], pd=100%, ROPE=[-26.42, 26.42], 0% in ROPE) suggesting that RSA leads to smaller lexicon sizes. An interaction effect between context-based and utility-based pragmatics (M=-41.42, CrI=[-58.62, -24.36], pd=100%, ROPE=[-26.42, 26.42], 2.00% in ROPE) that was probably significant, suggested that adding RSA reduced lexicon sizes mainly for context-aware trained agents but not for context-unaware trained agents. How did agents in the different conditions optimize the tradeoff between lexicon informativeness and size? In Fig 7.7A-D, we plot the tradeoff between lexicon informativeness and size for each condition. We calculated size-concept ratios by dividing the number of unique messages sent by the number of unique concepts in the dataset. In the context-unaware condition (Fig 7.7A), the messages
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 95 (a) Context-unaware (b) Context-unaware + RSA (c) Context-aware (d) Context-aware + RSA Figure 7.6: Distribution of message lengths for different levels of the conceptual hierarchy. Table 7.5: Lexicon sizes. Dataset # concepts Context-unaware + RSA Context-aware + RSA (3,4) 250 99.6 ±4.51 51.6 ±4.34 153.8 ±32.34 42.4 ±5.55 (3,8) 1460 517.2 ±10.13 379.8 ±28.55 902.4 ±116.91 155.4 ±14.86 (3,16) 100 24.6 ±9.4 56.4 ±9.99 40.0 ±17.39 39.0 ±9.77 (4,4) 1250 736.6 ±38.81 265.0 ±63.53 821.6 ±117.53 155.6 ±25.21 (4,8) 100 42.0 ±17.25 79.4 ±3.05 66.4 ±24.5 54.4 ±6.02 (5,4) 100 43.8 ±13.59 52.6 ±21.7 82.6 ±12.3 51.4 ±4.1 Mean number of unique messages and standard deviations over five runs are calculated for each dataset. that the agents sent were very informative, i.e. they singled out specific concepts. No more than six messages were needed to refer to ten concepts on average. The lexicon of context-aware agents was less informative than that (Fig 7.7C): Lexicon informativeness was much lower and ranged from 3.21 to 4.11 for the different datasets. In a Bayesian hierarchical model predicting lexicon informativeness by context-based pragmatics (context-unaware vs. context-aware), utility-based pragmatics (without RSA vs. with RSA) and their interaction, we found a main effect for context-based pragmatics (M=0.50, CrI=[0.39, 0.61], pd=100%, ROPE=[-0.11, 0.11], 0% in ROPE), a main effect for utility-based pragmatics (M=0.43, CrI=[0.32, 0.54], pd=100%, ROPE=[-0.11, 0.11], 0% in ROPE) and an interaction
96 7 Case study 3: The role of pragmatic mechanisms in language use and emergence effect (M=0.44, CrI=[0.33, 0.55], pd=100%, ROPE=[-0.11, 0.11], 0% in ROPE). These effects support the above observations. The mean lexicon size-concept ratios for each dataset were smaller in the context-unaware than in the context-aware condition. However, when adding RSA, this pattern was reversed (Fig 7.7B,D). In a Bayesian hierarchical model predicting lexicon size-concept ratio by context-based pragmatics (context-unaware vs. context-aware), utility-based pragmatics (without RSA vs. with RSA) and their interaction, we found a main effect for context-based pragmatics (M=-0.03, CrI=[-0.04, -0.01], pd=99.90%, ROPE=[-0.02, 0.02], 33.63% in ROPE) with undecided significance, a main effect for utility-based pragmatics (M=0.08, CrI=[0.06, 0.09], pd=100%, ROPE=[- 0.02, 0.02], 0% in ROPE) and an interaction effect (M=-0.09, CrI=[-0.10, -0.07], pd=100%, ROPE=[-0.02, 0.02], 0% in ROPE). These effects suggested that adding RSA improved the lexicon size-concept ratio for context-aware but not for context-unaware agents. (a) Context-unaware (b) Context-unaware + RSA (c) Context-aware (d) Context-aware + RSA Figure 7.7: Tradeoff between lexicon size and informativeness. The lexicon size is normalized by the number of concepts in a dataset. Lexicon-level efficiency of a language can also be quantified by how well the language follows a Zipf’s law like distribution of messages. The first part of Zipf’s law states that an efficient language contains few expressions which are used very frequently in a language (Zipf, 1949). In Fig 7.8A-B, we present the frequency distribution of the messages that have been sent in the four different game scenarios, context-aware and context-unaware with and without RSA. 5 Messages were ordered by their frequency rank on the x-axis and their relative frequency in the protocol is presented on the y-axis. We observed that in the context-unaware baseline, almost all messages were used with similar 5Frequency distributions for all datasets can be found in S5 Appendix: Frequency rank distributions for all datasets.
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 97 frequency. In the context-aware condition, some messages were used with very high frequency compared to the rest of the messages leading to the characteristic logarithmic relationship between relative frequency of messages in the protocol and their respective frequency rank. Adding RSA moved the respective distributions closer to a distribution resembling Zipf’s proposed relationship and natural languages, specifically the frequency distributions of English and Arabic corpora. The distribution of context-aware + RSA is closest to natural language. We conclude that the more we allowed for pragmatic mechanisms to be exploited in the simulations, the more the message distribution resembled a Zipf’s law like distribution of messages where there were few messages which were used very frequently and many messages which were used infrequently. (a) Context-unaware and context-aware (b) Context-unaware + RSA and context-aware + RSA Figure 7.8: Zipf’s law like distribution of message frequency plotted for dataset D(4,4). D(4,4) is a medium-sized dataset where objects consist of four attributes which can each take four different values. The second part of Zipf’s law states that languages maximize their efficiency by using shorter expressions more frequently than longer expressions (Zipf, 1935). In Fig 7.9A-B, we present the distribution of messages as a relation between frequency rank on the x-axis and message length on the y-axis. We observed again that the more we allowed for pragmatic mechanisms to be exploited in the simulations, the more the message distributions resembled Zipf’s law and natural languages with shorter messages being used more frequently than longer messages. (a) Context-unaware and context-aware (b) Context-unaware + RSA and context-aware + RSA Figure 7.9: Zipf’s law like distribution of message length plotted for dataset D(4,4). Lastly, we present a qualitative analysis of messages that were used in the four different test conditions, context-unaware, context-unaware + RSA, context-aware and context-aware + RSA. Table 7.6 shows examples for the randomly picked specific test concept (2,0,3) in dataset D(3,4), which is the smallest dataset where objects have three attributes that can each take four different values. We use a tuple notation for concepts, e.g. (2,0,3), where only fixed attributes are specified and attributes which are not fixed by a concept but can take any value are represented as “_”, e.g. (2,_,0). For a more intuitive interpretation of the results, we give natural language examples.
98 7 Case study 3: The role of pragmatic mechanisms in language use and emergence For example, (2,0,3) could mean small blue circle if the attributes were size, color and shape. These natural language examples are only possible interpretations of the concepts, the agents were trained on symbolic object vectors specifying only numerical attributes. We show examples for a coarse context, where the concept has been presented in a context where the distractor objects differed from the target concept in each attribute and a fine context, where the distractor objects in the context shared all but one attribute with the target concept. We list the messages that were sent and analyze their former use during training as well as which attributes of the concept the message probably communicated. In the context-unaware condition, the speakers are not aware of the context. This means that to be communicatively successful, they should communicate all attributes that are relevant to the target concept. Indeed, the message that was used to refer to (2,0,3) (small blue circle) during testing (“1,1,1,1”) was used to refer to similar specific concepts during training, such as (2,1,3) (small red circle) and (2,3,3) (small green circle), or to refer to related more generic concepts, i.e. (2,_,3) (small circle) and (2,0,_) (small blue). We observed no difference between the coarse and the fine condition in the context-unaware baseline. When adding RSA to the context-unaware baseline, the messages tended to be shorter and a more diverse set of messages was used. This was likely due to different aspects of the target concept being communcicated. We deduced this from the usage of the messages during training. By looking at the set of concepts that a message was used to refer to and their similarity, we reasoned about which attributes of the concept were likely communicated with this message. In the context-aware condition, we observed that the agents used different messages depending on the context condition (fine vs. coarse) and the specific distractor objects in the context. A likely strategy seemed to be that the agents compared the targets and distractors and communicated the one attribute that differed in the fine context condition or just one of the attributes that differed between targets and distractors in the coarse context condition. In the context-aware + RSA condition, agents used very similar messages to the messages that were used in the context-aware condition. The only difference was that they tended to use shorter messages where possible and, for example, reduced the messages “1,1,1,1” and “1,1,1” to “1,1”6. Discussion Overall, the results of Experiment 2 suggest that in novel communicative situations, the agents that had been trained in the different context conditions in Experiment 1, used different approaches to communicate the novel concepts and did so with varying degrees of success. The strategies of context-aware agents and context-aware + RSA agents lead to almost optimal communicative success, whereas context-unaware and context-unaware + RSA agents exhibited lower accuracies on the test dataset. These differences in performance were likely due to the different strategies that were used: Context-unaware speakers which had been trained without access to a shared context, came up with a strategy where they tended to communicate all fixed attributes of a target concept as predicted. This was shown by the qualitative analysis of messages and has implications 6 Sending the message “1” would in this case be even more efficient, but not possible in this scenario because the RSA agents are only able to use messages that have been produced during training at least once.
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 99 Table 7.6: Qualitative examples. Game scenario fixed indices fixed values context condition distractors communicated attributes messages use during training Contextunaware (1,1,1) (2,0,3) 0 (coarse context) (1,3,0) (2,0,3) [1,1,1,1] (2,1,3),(2,3,3) (0,1,0) (2,_,3), (2,0,_) (0,1,1) 2 (fine context) (0,0,3) (2,0,3) [1,1,1,1] see above (2,3,3) (2,1,3) Contextunaware + RSA (1,1,1) (2,0,3) 0 (1,3,0) (2,0,_) [1,2,1] (2,0,_) (0,1,0) (_,_,3) [1,1] (0,2,3),(0,_,3) (0,1,1) (_,_,3) [1,1] (2,3,3),(2,_,3) 2 (0,0,3) (2,_,_) [1,4,1,3] (2,1,_) (2,3,3) (_,0,_) [2] (1,0,3),(1,0,_) (2,1,3) (_,0,_) [2] (_,0,_) Contextaware (1,1,1) (2,0,3) 0 (1,3,0) (2,_,_) [4] (2,_,_) (0,1,0) (2,_,_) [4] (0,1,1) (_,0,_) [1] None 2 (0,0,3) (2,_,_) [4] (2,3,3) (_,0,_) [1,1,1,1] (_,0,_) (2,1,3) (_,0,_) [1,1,1] (_,0,_) Contextaware + RSA (1,1,1) (2,0,3) 0 (1,3,0) (2,_,_) [4] (2,_,_) (0,1,0) (2,_,_) [4] (0,1,1) (2,_,_) [4] 2 (0,0,3) (2,_,_) [4] (2,3,3) (_,0,_) [1,1] (_,0,_) (2,1,3) (_,0,_) [1,1] Qualitative examples for all four conditions and the dataset D(3,4) in response to the specific concept (2,0,3) (all attributes fixed). for the efficiency of the language used in this condition. Context-unaware agents used a smaller but very informative lexicon compared with the context-aware agents because they did not adapt their messages to the context conditions. Their messages tended to be longer, the more attributes were fixed in a target concept, suggesting that they communicated more attributes when the target concept was more specific. The messages used in the context-unaware condition did not follow Zipf’s law, i.e. frequently used messages were not shorter nor used in more communicative situations. When adding utility-based pragmatic abilities in the form of RSA, context-unaware speakers sent messages that were more efficient and shorter. This effect was more pronounced when agents had been trained context-unaware than when they had been trained context-aware. This was likely due to the context-unaware condition being the one in which messages were very long in the first place. Even though the messages produced by context-aware speakers were already almost optimal with regards to their message length, adding RSA still improved the efficiency of the messages chosen by context-aware speakers that reasoned about their interlocutor’s likely interpretation with RSA. That RSA improved the efficiency of messages is unsurprising given that the utility function
106 7 Case study 3: The role of pragmatic mechanisms in language use and emergence side. This means that the communication system can become simpler by being less informative because listeners can reason about the context and resolve the meaning of polysemous messages in context. In turn, this does not happen when speakers are trained without access to a (shared) context. Hence, our results provide evidence for the role of pragmatics in shaping category structure and communication about categories. We can consider the relationship between environment, communicative need and the category system to be modified by the availability of context. As predicted by Rosch et al. (1976) and Rosch (1978), the category system of a language can afford to be simpler when the meanings of categories are interpreted in context. In addition to the influence of the availability of context-based pragmatics on the emergence of a language, we investigated the role of utility-based pragmatics in novel communicative situations. We found that speakers choose more efficient messages when they reason about the utility of the message in terms of the listener’s likely interpretation via RSA-based reasoning. In summary, we have shown how context-based pragmatics shapes a category system when concepts at different levels of abstraction need to be communicated in interaction. Utility-based pragmatics has been shown to be especially effective when used for an emerging language that includes ambiguity in the context-aware condition, and its main function in our model is to make the communication about concepts at different levels of abstraction more efficient from the speaker perspective. Our findings suggest that efficient category systems that are reflected in an efficient linguistic system can emerge in agent-based models of communication in interaction. This has important implications for the question how human category and language systems might have evolved and are continuously shaped by the environment and a communicative need, namely the idea that pragmatics, in the form of context and utility-based reasoning, plays a crucial role in this process. 7.6 Methods 7.6.1 Formalization We use a multi-agent model of emergent communication. In emergent communication models, a speaker and a listener agent need to communicate about a target in a given context. In one round of the game, speaker and listener observe an input each, and then the speaker sends a message to the listener that describes the target. The listener interprets the message and selects the object they believe is the target, i.e. the one they assign the highest probability. During several iterations of training, the agents converge on a language-like system, i.e. a consistent mapping between targets and messages (Lazaridou et al., 2017,2018). To prevent agents from communicating about lower-level features that are not relevant from a human perspective, we use a concept-level reference game for training (Kobrock, Ohmer, et al., 2024; Mu & Goodman, 2021). This means that the agents do not communicate about a single object, but rather about a concept. We achieve this by creating concepts that consist of multiple target objects (Kobrock, Ohmer, et al., 2024; Mu & Goodman, 2021), see Fig 7.1.
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 107 The model is formalized as a communication game 𝐺=(𝑇𝑆, 𝐷𝑆, 𝑇𝐿, 𝐷𝐿) between a speaker agent 𝑆 and a listener agent 𝐿 . Both agents receive their own set of input objects, comprised of game size 𝑔 target objects 𝑇={𝑡1, ..., 𝑡𝑔} and 𝑔 distractor objects 𝐷={𝑑1, ..., 𝑑𝑔} . The speaker agent receives their own sets of targets 𝑇𝑆 and distractors 𝐷𝑆 ordered such that targets come first. 7 The receiver agent receives their own sets of targets 𝑇𝐿 and distractors 𝐷𝐿 , but they are shuffled and the listener does not know which objects are targets and which are distractors. This means they receive an input 𝑋𝐿={𝑥𝐿 1, ..., 𝑥𝐿 𝑖} , where 𝑖=2·𝑔 . The individual objects that comprise the target concept and the context are not necessarily shared between listeners and speakers, but they belong to the same target concept and context condition. For example, if the target concept is blue, then what 𝑇𝑆 and 𝑇𝐿 have in common is that all objects in these sets are blue. However, it does not matter, whether the target objects are blue circles, blue squares, blue triangles etc. These are randomly sampled from the set of blue objects. Similarly, in a context condition where one concept-defining attribute is allowed to be shared between 𝑇𝑆 and 𝐷𝑆 and between 𝑇𝐿 and 𝐷𝐿 , then the specific objects in the sets 𝐷𝑆 and 𝐷𝐿 , may differ. However, both distractors presented to the speaker and distractors presented to the listener differ in the same concept-defining attributes from the target concept. For example, blue circles are presented in a context comprised of circles of different colors (see Fig 7.2B). The listener’s task is to predict a label 𝑦𝐿 𝑖∈ {0,1} (0: distractor, 1: target) for each object 𝑥𝐿 𝑖 based on a message 𝑚 that it receives from the speaker. The speaker generates a message 𝑚 by choosing up to 𝑀 symbols from a vocabulary 𝑉 . The vocabulary is comprised of primitive discrete symbols, ranging from 0 to 𝑉 , i.e. "0", "1", "2", etc. (Lazaridou et al., 2018). The symbol 0 is defined as the end-of-sequence (EOS) symbol that can be used to terminate a message before 𝑀 , the maximum message length, is reached (Kobrock, Ohmer, et al., 2024; Mu & Goodman, 2021; Ohmer, Duda, & Bruni, 2022). All other symbols can be sent without such implications. They do not have a meaning in the beginning of training, but their meaning is negotiated and converged upon through training, i.e. through several iterations of playing the concept-level reference game. We train the speaker and listener agents with a joint binary cross entropy (BCE) loss maximizing the probability that the listener agent correctly identifies the target objects in their input: L 𝐵𝐶𝐸(𝑆, 𝐿, 𝐺)=−X 𝑖 log 𝑝𝐿(𝑦𝐿 𝑖|𝑥𝐿 𝑖,ˆ 𝑚),(7.2) where ˆ 𝑚∼𝑝𝑆(𝑚|𝑇𝑆, 𝐷𝑆)and 𝑝𝐿(𝑦𝐿 𝑖|𝑥𝐿 𝑖,ˆ 𝑚)=ReLu(GRU𝐿(ˆ 𝑚) · embed(𝑥𝐿 𝑖)). We implement the game using the EGG framework (Kharitonov et al., 2019) developed for implementing emergent communication games. Both our speaker and listener agents are implemented as neural networks with feed-forward (dense) layers for embedding the input objects and GatedRecurrent-Unit (GRU, Cho et al., 2014) cells to encode (speaker) or decode (listener) the messages. GRU cells are a specific type of Recurrent Neural Networks (RNNs) which have been proven useful for dealing with sequence data such as language. The GRUs consist of a single hidden layer and a gating mechanism that weighs how much previous information is considered when processing the current input. They are thus well-suited for detecting dependencies and patterns in language-like 7 If agents are trained context-unaware, the speaker agent does not process the distractor objects, but only the target objects.
108 7 Case study 3: The role of pragmatic mechanisms in language use and emergence input. Modeling a pragmatic speaker with the Rational Speech Acts (RSA) framework We model an RSA speaker following the Rational Speech Acts (RSA) framework (Frank & Goodman, 2012; Franke & Degen, 2016; Goodman & Frank, 2016). The RSA model provides a framework to model different listener and speaker types that vary by the information they take into account to choose or interpret a message. The base level speaker 𝑆0 chooses a message that describes the targets. The base level listener 𝐿0 interprets the message literally, i.e. it predicts which are the targets based on the learned meaning of a given message. The RSA speaker 𝑆RSA we use in this model is a level-1 speaker 𝑆1 , also called Gricean speaker (Franke & Degen, 2016), which maximizes relevant information as to allow the hearer to choose the correct target. This means that based on the assumption that a literal listener 𝐿0 chooses all targets of which the message is true with equal probability, the RSA speaker chooses a message that maximizes the probability of the listener to select the targets. For example, if the speaker needs to communicate a blue circle in a fine context consisting of circles of different colors than blue, then 𝑆0 would choose the messages “blue” or “circle” with equal probability. 𝐿0 would then choose targets of which the message is true, leading to communication success if the message was “blue” and to communication failure if the message was “circle”. An 𝑆1 speaker that we call 𝑆RSA in our model, would choose the message “blue” with higher probability than “circle” because they reason about the listener’s likely interpretation of the message. We implement 𝑆RSA as a speaker that maximizes a utility function following standard RSA models (Frank & Goodman, 2012; Franke & Degen, 2016; Goodman & Frank, 2016): 𝑆RSA(𝑇𝑆, 𝐷𝑆)=argmax 𝑚∈M𝑈(𝑚|𝑇𝑆, 𝐷𝑆),(7.3) where 𝑈 is the utility function, typically defined as 𝑈(𝑚|𝑤)=log 𝑝𝐿0(𝑤|𝑚) (Goodman & Frank, 2016). This means that speakers are rewarded if they choose a message that maximizes the loglikelihood of listeners interpreting the world 𝑤 correctly given message 𝑚 . Often this basic utility function is refined, for example by adding a cost term that penalizes long messagesScontras et al., 2021. In our implementation, we follow previous work (Fang et al., 2022) and define the utility function as: 𝑈(𝑚|𝑇𝑆, 𝐷𝑆)=log 𝑝𝐿(𝑦𝐿 𝑖|𝑥𝐿 𝑖, 𝑚) − 𝐶(𝑚),(7.4) where 𝐶(𝑚) is a cost applied to the message length and 𝑝𝐿(𝑦𝐿 𝑖|𝑥𝐿 𝑖, 𝑚) is the log-likelihood of the listener to select the correct targets, i.e. to predict the correct labels 𝑦𝑖 based on a message and input objects 𝑥𝑖=(𝑇𝑆, 𝐷𝑆) . The speaker evaluates the utility based on an internal listener model which is the trained agent from the first experiment 𝐿0 . We define log 𝑝𝐿(𝑦𝐿 𝑖|𝑥𝐿 𝑖, 𝑚) as the logits output from the internal listener model averaged for targets and distractors. To account for the fact that in typical RSA models, the messages that maximize the probability of choosing the correct targets at the same time minimize the probability of selecting the distractors, we calculate the overall utility as P𝑦𝑖=1 |𝑦𝑖=1|−P𝑦𝑖=0 |𝑦𝑖=0| . Importantly, the RSA speaker 𝑆RSA only has access to their targets 𝑇𝑆 and
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 109 distractors 𝐷𝑆 and thus is only able to infer how the listener would interpret a message referring to the target concept the speaker observes in a given context. This preserves some naturalness of the communicative scenario by preventing speakers from being omniscious. Finally, the accuracy is calculated on how the listener actually interprets a message in their own context 𝑥𝑖 consisting of 𝑇𝐿 and 𝐷𝐿 . The cost function 𝐶(𝑚)=𝜆|𝑚| penalizes messages based on their length, i.e. the number of symbols in a message. We use a cost factor of 1 for our simulations. 7.6.2 Dataset We construct six symbolic datasets that contain concepts with varying levels of specificity (ranging from specific to generic) depending on how many attributes are shared within a concept, i.e. between the different target objects. Objects are symbolic vectors comprised of 𝑛 attributes where each attribute can take 𝑘 values. For example, (1,2,1) is an object from the dataset D(3,4). We use the notation D( 𝑛 , 𝑘 ) to denote datasets of different sizes, ranging from three to five attributes which can each take between four and 16 values. Target objects that belong to a specific concept share all attributes. Target objects that belong to a generic concept share only one attribute. This means that the more specific a concept is, the more information needs to be communicated when describing the concept. Each concept is presented in a context. Depending on how many attributes are shared between the target objects and the distractor objects, we compare contexts ranging from fine contexts, in which all but one attributes are shared, to coarse contexts, in which one attribute is shared between targets and distractors. The context granularity also relates to how much information needs to be communicated to describe a given concept. As a general rule of thumb, the finer the context, the more information needs to be communicated. In Fig 7.10A-C, we present examples for concept and context combinations in the datasets. We use shapes with different colors and sizes only for visualization purposes. The agents are trained on symbolic vectors as described above. 60% of the concepts are assigned to the train split and 20% of the concepts are assigned to the validation split of the dataset. Each concept is presented in exactly one context, where the number of shared attributes is randomly sampled and shared between speaker and listener inputs. The data split used for testing (20% of the data) contains novel concepts in randomly sampled contexts. Figure 7.10: Dataset examples. The three objects in the top row are the target objects that form a target concept together. The three objects in the bottom row are the distractor objects that form the context. A: An example for a specific concept large blue circle in a fine context (two attributes shared). B: An example for a specific concept large blue circle in a coarse context (no attribute shared). C: An example for a generic concept circle in a coarse context.
110 7 Case study 3: The role of pragmatic mechanisms in language use and emergence 7.6.3 Hyperparameters and training We use the following hyperparameters: The agents have internal GRUs with a hidden size of 128 to encode and decode messages. We trained with a batch size of 16 and a learning rate of 0.001. The datasets were created with game size 10, i.e. 10 targets and 10 distractors per game round, and scaling factor 10, i.e. each concept is represented in the dataset 10 times. For training, we make use of the Gumbel-Softmax function that makes it possible to use backpropagation (Jang et al., 2017), a loss pressure that penalizes the length of messages and an early stopping criterion to make training more efficient. Details can be found in S1 Appendix: Training specifics and hyperparameters. 7.6.4 Evaluation We will evaluate the models and emerging languages according to the following criteria: First, we will report accuracies on the train, validation and test datasets. Accuracies on the train and validation datasets give us a measure of communicative success, i.e. how well do the listener agents perform, or, in other words, how often do they select the correct target objects. Accuracies on the test datasets additionally can be used as a measure for generalization abilities. Here, we can see whether the emerging language generalizes well to unseen concepts. Second, to investigate whether certain game scenarios lead to the emergence of more efficient languages, we need a measure that captures the efficiency of an emerging language. In the literature, efficiency has been related to small vocabulary sizes and to short messages lengths while communicating the same amount of information (Piantadosi et al., 2012). While the vocabulary size is in principle fixed in our simulations, we can measure the message lengths of an emergent language directly after training. We will also calculate the number of unique messages used to refer to concepts as an approximation to the agents’ lexicon size (Lazaridou et al., 2018). We will relate the size of the lexicon to the lexicon’s informativeness. The informativeness 𝐼 of a lexicon 𝐿 is defined as the average over the message informativeness 𝐼𝑚 over 𝑁 interactions following (Gualdoni & Boleda, 2024): 𝐼𝐿=1 𝑁 𝑁 X 𝑖=1 𝐼𝑖 𝑚 with 𝐼𝑚=1 𝑆𝑚, where 𝑆𝑚=1 𝑁X 𝑖 X 𝑗≠𝑖 𝑑(𝐶𝑖, 𝐶𝑗) is the spread of features that is calculated based on the average distance between concepts 𝐶 that have been referred to by 𝑚 . This means that messages receive a lower informativeness score if they are used to refer to concepts with higher distance and that messages receive a higher informativeness score if they are used to refer to concepts with lower distance, i.e. higher similarity. This captures the intuitive idea that words in a language which refer to a highly specific concept such as dalmatian are more informative than words which refer to a highly generic concept such as animal. A natural
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 111 language lexical system contains both highly specific and highly generic references and optimizes the tradeoff between size and specificity. Third, relating back to the proposed efficiency tradeoff between production and comprehension, we expect that an efficient language should contain ambiguity, i.e. synonymous and polysemous expressions. To measure whether an emergent communication protocol contains synonymous and polysemous messages, we calculate information-theoretic scores on the set of 𝐾 messages 𝑀={𝑚1, ..., 𝑚𝐾} uttered and the set of 𝐿 target concepts 𝐶={𝑐1, ..., 𝑐𝐿} the agents communicated about during the simulation. We calculate these scores on the final interactions between speakers and listeners in the last training epoch. Fig 7.11 shows the sets of concepts and messages (displayed as circles) with instances of the concepts and messages displayed as dots. The relevant information-theoretic scores are calculated on these two sets. A language that is maximally efficient only for the listener but not for the speaker should contain only one-to-one mappings between concepts and messages. The amount of one-to-one mappings in a language can be calculated with the normalized mutual information score between concepts and messages: 𝑁𝑀𝐼(𝐶, 𝑀)=𝐻(𝑀) − 𝐻(𝑀|𝐶) 0.5· (𝐻(𝐶) + 𝐻(𝑀)).(7.5) If this score is 1, then the efficiency tradeoff is maximal for the listener and minimal for the speaker. If this score is 0, on the other hand, we cannot conclude that efficiency for the speaker is high; as it might be that a completely random mapping between concepts and messages have been learned that would also be highly inefficient for the speaker. Thus, to address efficiency for the speaker, we consider how much ambiguity, i.e. how many synonymous and polysemous messages, an emergent protocol includes. To do this, we calculate scores based on conditional entropies. The conditional entropy of messages given concepts, 𝐻(𝑀|𝐶)=−X 𝑐∈𝐶,𝑚∈𝑀 𝑝(𝑐, 𝑚)log 𝑝(𝑐, 𝑚) 𝑝(𝑐),(7.6) measures how much uncertainty remains about the messages after knowing the concepts, i.e. the signal uncertainty. High signal uncertainty means that there is a one-to-many mapping of meanings-to-signals (Winters et al., 2015). In other words, when knowing the concept, it is highly uncertain which message has been used to refer to it. This measure captures the synonymy of messages (Ohmer, Duda, & Bruni, 2022; Winters et al., 2015). The consistency score uses the conditional entropy 𝐻(𝑀|𝐶) to measure how much uncertainty remains about the messages after knowing the concepts. It is calculated as follows: consistency(𝐶, 𝑀)=1−𝐻(𝑀|𝐶) 𝐻(𝑀).(7.7)
112 7 Case study 3: The role of pragmatic mechanisms in language use and emergence Conversely, the conditional entropy of concepts given messages, 𝐻(𝐶|𝑀)=−X 𝑐∈𝐶,𝑚∈𝑀 𝑝(𝑚, 𝑐)log 𝑝(𝑚, 𝑐) 𝑝(𝑚),(7.8) measures how much uncertainty remains about the concepts after knowing the messages, i.e. the meaning uncertainty. High meaning uncertainty means that there is a one-to-many mapping of signals-to-meanings (Winters et al., 2015). In other words, when knowing the message, it is highly uncertain which concept it refers to. This measure captures the polysemy of messages (Ohmer, Duda, & Bruni, 2022; Winters et al., 2015). The effectiveness score uses the conditional entropy of concepts given messages 𝐻(𝐶|𝑀) to measure how much uncertainty remains about the concepts after knowing the messages. It is calculated as follows: effectiveness(𝐶, 𝑀)=1−𝐻(𝐶|𝑀) 𝐻(𝐶).(7.9) If the effectiveness score is maximal, i.e. 1.0, this means that the language is highly effective, or, in other words, that there are no polysemous expressions in the lexicon. Figure 7.11: Information-theoretic scores and ambiguity in language. The set of concepts 𝐶 is the red circle on the left (red and purple), the set of messages 𝑀 is the full blue circle on the right (blue and purple). The mutual information 𝐼(𝐶, 𝑀) (only purple) captures one-to-one mappings between messages and concepts. The conditional entropy 𝐻(𝑀|𝐶) (only red) is related to synonymous mappings between messages and concepts. The conditional entropy 𝐻(𝐶|𝑀) (only blue) is related to polysemous mappings between messages and concepts. Fourth, as another measure of efficiency and to assess whether the emerging languages share efficiency related properties with human languages, we analyze the frequency distributions of messages. We investigate the following two relationships proposed by Zipf’s law (Zipf, 1935; Zipf, 1949) which seem to be a common principle of natural languages on the lexicon level (Cancho & Solé, 2003; Chaabouni et al., 2019; Piantadosi et al., 2011). First, when ordering messages according to their frequency, i.e. their rank in a lexicon, it has been shown that the relative frequency of a word in a corpus decays exponentially with each rank. This means that few messages are used very frequently (Cancho & Solé, 2003). Second, word lengths are determined by their frequency,
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 113 where the most frequent words tend to be the shortest in a language (Piantadosi et al., 2011). We will test these two properties on the agent’s protocols of messages used during inference on novel test concepts and compare the resulting distributions in the different conditions and to natural language data from English and Arabic. We use the Leipzig Corpora Collection (Goldhahn et al., 2012) and specifically the News copora with 10K words for English (2024) and Arabic (2022) including words and their frequencies in the corpus. Frequency ranks are determined by sorting the words/messages by their frequency of occurence 𝐹={𝑓1, ... 𝑓𝑚} . The relative frequency 𝐹𝑟𝑒𝑙 of a word is calculated as 𝐹𝑟𝑒𝑙 =𝑓𝑖 P𝑚 𝑗𝑓𝑗 , where 𝑚 is the number of words/messages and 𝑖 is the 𝑖 th word. We choose 𝑚=30 for both artificial simulations and natural languages, i.e. we calculate the frequency distributions for the 30 most frequent words/messages. We calculate message lengths with a Python script and the PyArabic library (Zerrouki, 2023). Acknowledgments The simulations were run on a high-performance computing cluster funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 456666331. Kristina Kobrock was supported by the DFG-funded Research Training Group “Computational Cognition” (DFG-GRK 2340). We thank Muhip Tezcan, Eosandra Grund, and Isabella del Pozo for their assistance in implementing parts of the software used in this project. Author Contributions: Kristina Kobrock: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Writing - Original Draft, Writing - Review & Editing, Visualization. Xenia Ohmer: Conceptualization, Methodology, Software, Writing - Review & Editing. Elia Bruni: Conceptualization, Methodology, Writing - Review & Editing, Supervision. Nicole Gotzner: Conceptualization, Methodology, Writing - Review & Editing, Supervision. Appendix A S1 Appendix: Training specifics and hyperparameters. Gumbel-Softmax relaxation We train with the Gumbel-Softmax relaxation that makes it possible to use backpropagation (Jang et al., 2017). We use a temperature of 2.0 and a temperature update of 0.99.
114 7 Case study 3: The role of pragmatic mechanisms in language use and emergence Length cost pressure We train the agents with a loss pressure that penalizes the length of messages by multiplying the symbol’s position in a message with a cost factor 𝑐𝑓 . We choose the cost factor 𝑐𝑓=0.001 . We set the maximum message length 𝑀 to 20 symbols and the vocabulary size 𝑉 to the number of values in a dataset plus one for the end-of-sequence (EOS) symbol. There are two main reasons for applying such a length pressure during training. The first reason is that the length cost pressure works like an efficiency pressure on the emerging language. If we expect differences between the context-unaware and the context-aware scenario, then these should be due to efficiency reasons. While human communication is typically costly, with neural network agents sending bits of information is cheap and we thus need to include a pressure that incentivizes the agents to use efficient language. The second reason is that for the RSA simulations to work, we want the training to generate an emerging language that contains messages of different lengths. These can be better exploited by the RSA senders. If messages generated during training were all of the same length, then the cost term in the RSA utility function would have no effect. Early stopping We use early stopping for training. We stop training when a validation accuracy of 0.90 has been reached at least once during training and the validation loss has not changed more than 0.001 during 10 epochs. In case these criteria are not met during 300 epochs of training, we stop nevertheless. B S2 Appendix: Statistical models In Tables 7.7 and 7.8, we specify the model syntax and priors used for the statistical models fitted for Experiment 1 and 2, respectively. Table 7.9 presents the posterior summaries for the Bayesian hierarchical models predicting NMI, effectiveness and consistency for Experiment 1. Table 7.7: Statistical models fitted with brm for Experiment 1. Model Syntax Priors 1) Accuracy validation accuracy ∼ context + (1|dataset) Intercept : uniform(0,1) , sigma: uniform(0,0.1) 2) Message length message length ∼context ∗ fixed attributes + (1|dataset) Intercept : uniform(0,20) , sigma: uniform(0,5) 3) Entropy NMI ∼context + (1|dataset)Intercept : uniform(0,1) , sigma: uniform(0,0.1) effectiveness ∼context + (1|dataset) consistency ∼context + (1|dataset) 4) Entropy * concept hierarchy NMI ∼context ∗ fixed attributes + (1|dataset) Intercept : uniform(0,1) , sigma: uniform(0,0.1)effectiveness ∼context ∗ fixed attributes + (1|dataset) consistency ∼context ∗ fixed attributes + (1|dataset)
7 Case study 3: The role of pragmatic mechanisms in language use and emergence 115 Table 7.8: Statistical models fitted with brm for Experiment 2. Model Syntax Priors 1) Accuracy test accuracy ∼context ∗ RSA + (1|dataset) Intercept : uniform(0,1) , sigma: uniform(0,0.1) 2) Message length message length ∼context ∗ RSA ∗fixed attributes + (1|dataset) Intercept : uniform(0,20) , sigma: uniform(0,5) 3) Lexicon size lexicon size ∼context∗RSA + (1|dataset) Intercept : uniform(0,1460) , sigma: uniform(0,100) 4) Lexicon informativeness lexicon info ∼context∗RSA+ (1|dataset) Intercept : uniform(0,9) , sigma: uniform(0,2) 5) Lexicon size-concept ratio lexicon size-concept ratio ∼ context ∗RSA + (1|dataset) Intercept : uniform(0,1) , sigma: uniform(0,0.1)
122 7 Case study 3: The role of pragmatic mechanisms in language use and emergence Regarding the role of pragmatics in this joint evolution and in the use of a language for concrete references to concepts, we compared the availability of context-based and utility-based pragmatics to a baseline where no pragmatic mechanisms were available. Considering the role of contextbased pragmatics, we showed that languages tend to be more ambiguous and efficient when they emerged during interactions in a shared context. The speaker agents came up with a very efficient strategy where they made use of the context to communicate only the attributes that were needed to discriminate the target concept from the context. This led to a polysemous mapping between concepts and messages, but the ambiguity could be resolved in context (as was also proposed in Piantadosi et al., 2012). In language use, we again found increased efficiency at the message and lexicon levels when a shared context was available during the emergence of the language. Considering utility-based pragmatics, when speakers took the expected utility of an utterance into account, this improved the generalization accuracy and efficiency of the communication system. While the message-level efficiency can be attributed to the cost term in the RSA function, the improved lexicon-level efficiency is an emergent feature. This likely relates to the fact that RSA speakers can condense longer messages with the same meaning into a single message. When comparing agents equipped with different combinations of contextand utility-based pragmatic abilities, we found that contextand utility-based pragmatic mechanisms yield very similar results. Specifically, we observed that utility-based pragmatics works better when combined with a language that emerged when a shared context was available. This might mean that utility-based pragmatics and context-based pragmatics rely on similar mechanisms. One hypothesis might be that pragmatically efficient languages emerge in interactions where a shared context is available, leading to the conventionalization of terms where much of the context information is shared by both interlocutors and there is little uncertainty involved. We can further hypothesize that utility-based pragmatics is needed only when the conventionalized terms fall short of the communicative intentions. For example, if only the terms green and circle are conventionalized, but the concept to be communicated is a blue circle that should be discriminated from a green circle, then RSA-style recursive pragmatic reasoning can help to choose the utterance with highest utility, circle, in this context. In conclusion, we modeled the pressures for simplicity and informativeness as a tradeoff between the speaker’s and the listener’s needs in interaction. The communicative and cognitive need for efficiency likely plays a key role in optimizing this tradeoff and, as a result, in shaping linguistic and category systems. Our modeling results provide evidence for this hypothesis and show that the emergence of efficient category and linguistic systems can be a result of efficient interaction. Moreover, we provided evidence for the key role of pragmatics in shaping linguistic and category systems. It is widely accepted that pragmatics plays a key role in concrete reference situations, and also in shaping communicative systems. Here, we argue that pragmatics also plays a key role in shaping category systems due to the same mechanisms and the proposed joint evolution of category and linguistic systems. This is a key finding for the view that language and cognition interact.
8 Case study 4: Linguistic strategies for generalization and abstraction This chapter presents case study 4. The chapter starts with a high-level introduction followed by the content of the publication: Kobrock, K., Ohmer, X., Bruni, E., & Gotzner, N. (2025a). Agents generalize to novel levels of abstraction by using adaptive linguistic strategies. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Findings of the Association for Computational Linguistics: ACL 2025 (pp. 8685– 8699). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findingsacl.455 The chapter ends with a brief summary of the main contributions of the publication and a discussion of its implications for the broader research question of this dissertation. 8.1 High-level introduction The goal of this case study is to investigate generalization and abstraction and how language can help achieve these abilities. For this purpose, we again use an emergent communication model similar to the one used in the previous two case studies. Here, the focus is on a zero-shot generalization task where agents need to generalize their language and refer to unseen concepts at novel levels of abstraction. We operationalize concepts again as a combination of objects that share a certain number of attributes. The more attributes shared within members of the same target concept, the more specific the concept is. The fewer attributes shared within members of a target concept, the more generic the concept is. We study the zero-shot generalization to novel levels of abstraction by manipulating the training and test data splits. In the “to generic” condition, we train agent pairs to communicate about more specific concepts, and test them on the most generic concepts. In the “to specific” condition, we train speaker and listener agents to communicate about more generic concepts, and test them on the most specific concepts. First, we compare the generalization performance of the agents in both conditions. Next, we investigate in detail which messages the agents use to generalize to concepts at novel levels of abstraction. We gain insights into the agents’ linguistic strategies by comparing the set of messages that agents use to generalize to the unseen concepts with the set of messages that were established during training. Specifically, we ask what the ratio of novel messages is and how many messages agents reuse from training. This helps us to gain insight into whether the agents spontaneously use novel labels to refer to the novel concepts. If agents indeed use novel labels, then the next question is, how these novel messages are composed. We hypothesize that in the “to specific” generalization condition, agents might produce a large amount of novel messages by compositionally combining symbols from previously established messages. In the “to generic” condition, however, we hypothesize that a compositional strategy will be less useful. We use qualitative analyses and quantitative metrics to gain a detailed understanding of the types of messages that the agents send in the different conditions and how these messages differ from those sent during training. This enables us to identify linguistic strategies
124 8 Case study 4: Linguistic strategies for generalization and abstraction for generalization to more specific concepts in the “to specific” condition, as well as for abstraction to more generic concepts in the “to generic” condition.
8 Case study 4: Linguistic strategies for generalization and abstraction 125 8.2 Abstract We study abstraction in an emergent communication paradigm. In emergent communication, two artificial neural network agents develop a language while solving a communicative task. In this study, the agents play a concept-level reference game. This means that the speaker agent has to describe a concept to a listener agent, who has to pick the correct target objects that satisfy the concept. Concepts consist of multiple objects and can be either more specific, i.e. the target objects share many attributes, or more generic, i.e. the target objects share fewer attributes. We tested two directions of zero-shot generalization to novel levels of abstraction: When generalizing from more generic to very specific concepts, agents utilized a compositional strategy. When generalizing from more specific to very generic concepts, agents utilized a more flexible linguistic strategy that involves reusing many messages from training. Our results provide evidence that neural network agents can learn robust concepts based on which they can generalize using adaptive linguistic strategies. We discuss how this research provides new hypotheses on abstraction and informs linguistic theories on efficient communication. 8.3 Introduction One of the most fundamental goals of Artificial Intelligence (AI) and Natural Language Processing (NLP) research is to build models which can generalize well to unseen data. This is, after all, one of the crucial abilities observed in human intelligence. Abstraction has been argued to be a necessary first step towards achieving generalization (Yee, 2019). But there are also alternative views such as the exemplar-based model of categories where generalization is achieved without abstraction (Ambridge, 2020; Daelemans, 2008). We believe that understanding abstraction and how it interacts with generalization, is fundamental to building well-generalizing models. Humans naturally use abstraction to solve complex tasks and to communicate about strategies and solutions. Well-designed AI and NLP systems cannot only benefit from good abstraction abilities in, for example, reasoning and solving complex tasks (Ho et al., 2019; Zheng et al., 2024), but interactive systems should also be able to deal with human language inputs which involve abstractions (e.g., Lachmy et al., 2022). Many researchers studying human abstraction argue for a role of language therein (see e.g., Gentner & Asmuth, 2019; Lupyan & Lewis, 2019; Sloutsky & Deng, 2019; Yee, 2019). The main idea of these accounts is that the lexicalization of concepts, i.e. having a label for a concept, helps to acquire and structure information we obtain about an entity and to observe commonalities within members of a concept in the first place. The role of language in abstraction can also be tested in computational systems. The goals of the current research are to understand how abstraction and generalization interact and ultimately to inform the improvement of AI and NLP systems towards achieving human-like abstraction and generalization abilities. Starting from the assumption that language is useful for abstraction, we study abstraction in a communicative setting and investigate how abstraction is achieved with the help of linguistic
126 8 Case study 4: Linguistic strategies for generalization and abstraction strategies such as compositionality and the reuse of previously established messages. To gain insights into the principled mechanisms of abstraction and the role of language for abstraction, we use a language emergence scenario. In language emergence research, the idea is to define a set of assumptions and then observe how these assumptions change a language-like system that emerges during interaction (see e.g., Chaabouni et al., 2020; Galke et al., 2022; Lazaridou et al., 2017; Rodríguez Luna et al., 2020). This modeling framework is ideal for investigating whether human-like behavior and communication strategies can emerge even in a comparatively simple communicative setup between two artificial neural network agents. In our model, two artificial neural network agents solve a reference game, where a speaker agent has to communicate a target concept to a listener agent who needs to select the correct target concept in a context. We operationalize a concept as a set of target objects following previous work (Kobrock, Ohmer, et al., 2024; Kobrock, Uhlemann, & Gotzner, 2024; Mu & Goodman, 2021). Our target concepts are designed in a hierarchical fashion, ranging from very specific concepts consisting of objects where all attributes (e.g., size, color and shape) are fixed to a certain value, e.g., ‘small blue circle’, to very generic concepts consisting of objects where only one attribute is fixed, e.g., ‘circle’. We can study abstraction by making use of this abstraction hierarchy. Here, we are interested in a specific kind of abstraction, namely the zero-shot generalization to concepts at novel levels of the concept hierarchy, or, to concepts at novel levels of abstraction (following the terminology of seminal research from Cognitive Psychology by Rosch et al., 1976). We will not only look at the generalization performance of the trained models, but also at the linguistic strategies the agents employ. Specifically, we investigate the properties of the emergent protocol and the use of novel vs. established messages during abstraction. While previous work in emergent communication has highlighted the role of compositionality for generalization (see e.g., Hazra et al., 2021; Kottur et al., 2017; Lazaridou et al., 2018), in our experiments we disentangle two directions of generalization and propose that they require different linguistic strategies. We find that agents use a compositional strategy only when generalizing to specific concepts, but not when generalizing to generic concepts. These results highlight that compositionality is not the only way to achieve generalization, which is in line with recent findings from Chaabouni et al. (2020) and Kharitonov and Baroni (2020). 8.4 Method 8.4.1 General setup We use an emergent communication paradigm (e.g., Chaabouni et al., 2019; Lazaridou et al., 2018) and build on the concept-level reference game developed in previous work (Kobrock, Ohmer, et al., 2024; Mu & Goodman, 2021). We train two artificial neural network agents, one speaker and one listener agent. Over several iterations, these agents develop a communication system by solving the following task: The speaker agent 𝑆 has to communicate a concept, i.e. a set of target objects
8 Case study 4: Linguistic strategies for generalization and abstraction 127 𝑇={𝑡1, ..., 𝑡𝑔} , to the listener agent 𝐿 whose task is to identify the correct targets among a set of distractors 𝐷={𝑑1, ..., 𝑑𝑔} . We call the set of target objects the concept and the set of distractor objects the context. The listener’s task is to identify the target concept in a certain context given a message generated by the speaker. The message is a vector of symbols generated by the speaker neural network which does not have a pre-specified meaning. Rather, the meaning of a message emerges over several interactions between the agents and is defined by its usage (see e.g., Lazaridou et al., 2017). Concepts vary in specificity, ranging from specific, where all attributes are shared among the target objects, to generic, where only one attribute is shared among the targets. Contexts can range from being fine, where all but one attributes are shared between targets and distractors, to being coarse, where no attribute is shared between targets and distractors. Both agent networks are trained in a Reinforcement Learning paradigm with the Gumbel-Softmax relaxation (Jang et al., 2017) on a joint loss that depends on whether the listener correctly identifies the targets and distractors given the speaker-generated message. 8.4.2 Zero-shot conditions and hypotheses We test the zero-shot generalization abilities of the trained networks in two conditions (see Figure 8.1): The first condition, “to specific”, tests whether agents are able to generalize to the most specific concepts when having seen more generic concepts during training. In this condition, we expect the emerging communication system to encode more generic concepts (such as “blue” or “circle”). For a successful zero-shot generalization, these more generic concepts would need to be combined to describe a specific concept (such as “blue circle”). Here, agents will need to combine previously learned attributes compositionally to describe a more specific concept. The second condition, “to generic”, tests whether agents are able to generalize to the most generic concepts when having seen more specific concepts during training. In this condition, we expect the emerging communication system to encode more specific concepts (such as “blue circle” or “orange circle”). For a successful zero-shot generalization, agents will need to abstract away from contextually irrelevant features and find the common attribute that all targets share (e.g., “circle”). 8.4.3 Dataset The agents are trained on six symbolic datasets developed in previous work (Kobrock, Ohmer, et al., 2024). These datasets contain all possible concepts, ranging from specific to generic, and contexts, ranging from fine to coarse, for a given number of attributes and values. For example, dataset D(3,4) contains all possible concepts and contexts given that objects in this dataset have three attributes and each attribute can take four different values. If we think of the three attributes as shape, color and size, an example for a specific concept would be “small blue circle” and an example for a generic concept would be “square”. In a fine context, objects belonging to the concept “small blue circle” would need to be discriminated against objects that are also small and blue. In a coarse context, distractor objects do not share any attributes with the target concept. This also means that there are
128 8 Case study 4: Linguistic strategies for generalization and abstraction … training testing … training testing 1) to specific: 2) to generic: Figure 8.1: Examples for speaker inputs for training and testing in the two zero-shot test conditions “to specific” and “to generic”. Each input consists of targets (i.e., concepts) in the green bounding box and distractors (i.e., context). more possible contexts for specific concepts than for generic concepts and the datasets reflect this relationship. We use a scaling factor of 10 to construct the datasets, i.e. each concept is included in a dataset 10 times.1This ensures that the datasets contain enough training data. For the zero-shot dataset generation, we manipulate the training, validation and test splits of the data. In the “to specific” condition, the test split contains all most specific concepts available, i.e. those where all attributes are shared among the targets. The training and validation splits are composed of the remaining concepts which are more generic with 75% of the data used for training and 25% of the data used for validation. In the “to generic” condition, the test split contains all most generic concepts available, i.e. those where only one attribute is shared among the targets. The training and validation sets contain the remaining more specific concepts with 75% of the data used for training and 25% of the data used for validation. Dataset sizes can be inspected in Tables 8.10 and 8.11 in Appendix Dand are comparable between zero-shot conditions. 8.4.4 Architecture and training A communication game between a speaker 𝑆 and a listener 𝐿 is defined as 𝐺=(𝑇𝑆, 𝐷𝑆, 𝑇𝐿, 𝐷𝐿) , where 𝑇𝑆={𝑡𝑆 1, ..., 𝑡𝑆 𝑔} and 𝐷𝑆={𝑑𝑆 1, ..., 𝑑𝑆 𝑔} are the inputs to the speaker, i.e. sets of game size 𝑔 targets and distractors, and 𝑇𝐿 and 𝐷𝐿 are the analogously defined inputs to the listener. For these inputs, 𝑇𝑆≠𝑇𝐿 and 𝐷𝑆≠𝐷𝐿 hold, i.e. the targets and distractors presented to the speaker differ from the targets and distractors presented to the listener to ensure communication of higher-level concepts (Kobrock, Ohmer, et al., 2024; Mu & Goodman, 2021). In each round of 1 We use this scaling factor only to construct the train and validation dataset splits. The zero-shot test is performed on a test split that contains the novel concepts only once.
8 Case study 4: Linguistic strategies for generalization and abstraction 129 the game, 𝑆 generates a message 𝑚=(𝑠𝑗)𝑗≤𝑀 , where 𝑠𝑗 is a symbol from vocabulary 𝑉 and 𝑀 is the maximal message length 2 , based on the inputs 𝑇𝑆 and 𝐷𝑆 . 𝐿 in turn, receives 𝑚 and an input 𝑋𝐿={𝑥𝐿 1, ..., 𝑥𝐿 𝑖} , where 𝑖=2·𝑔 which contains the targets 𝑇𝐿 and distractors 𝐷𝐿 shuffled. 𝐿 then predicts a label 𝑦𝐿 𝑖∈ {0,1} (0: distractor, 1: target) for each object 𝑥𝐿 𝑖 in its input (see e.g., Kobrock, Ohmer, et al., 2024; Mu & Goodman, 2021; Ohmer, Duda, & Bruni, 2022). We visualize the setup in Figure 8.2. Speaker neural network (GRU) Listener neural network (GRU) message [4,11,7,0] prediction and training Figure 8.2: Architecture: Speaker and listener neural networks receive separate inputs where target objects satisfy the same target concept (here “blue”) and distractor objects (i.e., the context) share the same number of attributes with the target concept (here 0). They are trained on successful communication, i.e. when the listener identifies the correct target objects. Fortheimplementation 3 ,weusetheEGGframeworkforemergentcommunicationgames(Kharitonov et al., 2019, MIT license). Both agents are implemented in a similar fashion: Feed-forward layers with 64 units serve as embedding layers for the input objects. The speaker targets and distractors are embedded separately and then concatenated into a joint embedding. The listener input objects are processed by just one embedding layer. For message encoding and decoding, both speaker and listener networks use single-layer Gated Recurrent Units (GRU, Cho et al., 2014) with a hidden layer size of 128 that can deal with sequential inputs of varying lengths. A speaker-listener pair is trained with binary cross entropy loss L 𝐵𝐶𝐸(𝑆, 𝐿, 𝐺)=−X 𝑖 log 𝑝𝐿(𝑦𝐿 𝑖|𝑥𝐿 𝑖,ˆ 𝑚),(8.1) where ˆ 𝑚∼𝑝𝑆(𝑚|𝑇𝑆, 𝐷𝑆) and 𝑝𝐿(𝑦𝐿 𝑖|𝑥𝐿 𝑖,ˆ 𝑚)=ReLU(GRU𝐿(ˆ 𝑚) · embed(𝑥𝐿 𝑖)) maximizing the probability that the listener correctly identifies targets and distractors with a label 𝑦𝑖∈ {0,1} (0: distractor, 1: target) for each object 𝑥𝑖 . To ensure differentiability for backpropagation, we use the straight-through Gumbel-Softmax trick (Jang et al., 2017) with temperature 𝜏=2 and a decay rate of 0.99. These and other hyperparameters were determined in a grid search that we conducted for all parameters over the different dataset sizes aiming for maximal validation accuracy. We train with batch size 32 and learning rate 0.001. For our simulations, we use game size 10, i.e. 10 target 2The end-of-sequence symbol 0can be used to terminate a message before 𝑀is reached. 3All code and analysis scripts are available at https://github.com/kristinakobrock/zero-shot-abstraction.
130 8 Case study 4: Linguistic strategies for generalization and abstraction objects form a concept and 10 distractor objects form the context. The maximum message length 𝑀 is defined as the total number of attributes in a dataset plus the End of Sequence (EOS) symbol 0 . The vocabulary size for each dataset corresponds to the total number of attribute values present. We establish a minimal vocabulary size for each dataset as the sum of the number of attribute values plus one additional symbol. This minimal vocabulary size is then scaled by a factor of 𝑓=3 , as suggested by Ohmer, Duda, and Bruni (2022) to ensure a sufficiently large communication channel (Chaabouni et al., 2020). 8.5 Results We trained the models on six symbolic datasets with varying numbers of attributes and values. In a dataset 𝐷(𝑛, 𝑘) , objects have 𝑛 attributes which each can take 𝑘 different values. For all metrics, we report means and standard deviations over five individual runs per dataset. 8.5.1 Generalization performance We evaluate the agents’ performance on the test datasets to assess their zero-shot generalization abilities. 4 Accuracies are calculated as a percentage over the objects that the listener classifies as targets or distractors. An accuracy of 0.9 means that 90% of the objects, i.e. 18 objects with a game size of 10, have been classified correctly as targets or distractors. Or, in other words, two objects have been misclassified. Table 8.1 summarizes the mean test accuracies over the five runs conducted on each dataset for both conditions. All zero-shot test accuracies are >=0.63 indicating that the listeners correctly identify more than 60% of the 20 objects as targets or distractors. This corresponds to a number of 12 correctly identified objects. Agents achieve higher performance in the “to specific” condition compared to the “to generic” condition in all datasets. Comparing test accuracies between datasets, generalization performance is better for datasets with more attributes. Specifically, on datasets with at least four attributes, agents achieve generalization accuracies of 0.82 or higher in both conditions. This means that speakers choose expressions to describe the held-out concepts at novel levels of abstraction that enable listeners to classify at least 16 of the 20 objects correctly. 4 Training and validation accuracies for both conditions are >=0.97 indicating that the agents have learned the task and achieved high performance on both the training and the validation data splits - a necessary prerequisite for a valid interpretation of the zero-shot test accuracies (see Tables 8.4 and 8.5 in Appendix A).
8 Case study 4: Linguistic strategies for generalization and abstraction 131 Table 8.1: Zero-shot test accuracies for both conditions. to specific to generic D(3,4) 0.92 ±0.02 0.71 ±0.04 D(3,8) 0.85 ±0.01 0.68 ±0.07 D(3,16) 0.82 ±0.03 0.63 ±0.03 D(4,4) 0.95 ±0.00 0.82 ±0.02 D(4,8) 0.95 ±0.01 0.82 ±0.07 D(5,4) 0.96 ±0.01 0.84 ±0.06 8.5.2 Concept reference We investigate the emergent mappings between concepts and messages during training with the Normalized Mutual Information (NMI) score calculated over messages 𝑀and concepts 𝐶: NMI(𝐶, 𝑀)=𝐻(𝑀) − 𝐻(𝑀|𝐶) 0.5· (𝐻(𝐶) + 𝐻(𝑀)) ,(8.2) The NMI score is maximal (i.e., 1.0) if for all messages and concepts seen during training, every message maps to exactly one concept and vice versa. In other words, a maximal score indicates that the agents developed a protocol that includes only one-to-one mappings between messages and concepts, i.e. no ambiguity. We expect high but not maximal NMI scores which would indicate that the agents have learned a structured but not unambiguous mapping between concepts and messages. The mean NMI scores calculated for messages and concepts during training in five runs range between 0.84 and 0.95 in the “to specific” condition, i.e. when trained on more generic concepts, and between 0.77 and 0.87 in the “to generic” condition, i.e. when trained on more specific concepts (see Table 8.2). This indicates that a structured communication protocol has emerged in both conditions, while more ambiguity arises when training the agents on more specific concepts in the “to generic” condition. Table 8.2: NMI scores for both conditions. to specific to generic D(3,4) 0.93 ±0.03 0.87 ±0.04 D(3,8) 0.95 ±0.01 0.82 ±0.02 D(3,16) 0.87 ±0.01 0.77 ±0.02 D(4,4) 0.94 ±0.01 0.87 ±0.05 D(4,8) 0.84 ±0.03 0.83 ±0.03 D(5,4) 0.87 ±0.02 0.83 ±0.04 8.5.3 Generalization strategies When agents generalize to novel concepts in the zero-shot test, there are two conceivable strategies. Firstly, agents might reuse messages that have been successfully used during training also on the test dataset. Secondly, agents might invent novel messages to describe the novel concepts in the