scieee AI-readable full text Open interactive document viewer

Instruction tuning modulates discourse biases in language models

Kankowski, Florian; Solstad, Torgrim; Zarrieß, Sina; Bott, Oliver

Abstract

Instruction tuning (IT) has been a fruitful technique for aligning Large Language Models with human preferences. However, the linguistic implications of IT remain unclear. In two experiments on coreference and coherence biases in the context of Implicit Causality, we investigate how IT modulates these discourse biases in relation to model size. Our results show that IT interacts with model size – instruction-tuned models display enhanced coherence biases and more human-like coreference patterns, sometimes exceeding human performance. However, this effect appears size-dependent, suggesting that IT causes some linguistic patterns to emerge that are dormant in the respective foundation models.

Full text

•Research question: Do model size and Instruction Tuning (IC) affect the emergence of Implicit Causality (IC) biases in LLMs? •Implicit causality (IC) is a discourse bias phenomenon involving predictions. • Widely researched in psycholinguistics1-5 • A number of studies of IC in LLMs. Some6-9 found no, or weak, effects of IC, others10 found LLMs to overshoot compared to human biases •Gaps in model size and model type: No, or weak effects for older / smaller models, overshooting in newer models with Instruction Tuning (ChatGPT 3.5, Vicuna) •Discourse ability gap: Existing studies have focused on coreference. However, coreference and coherence are intertwined.11 ØInteresting testbed for comparisons of predictive capabilities of humans12 and LLMs Florian Torgrim Sina Oliver Kankowski Solstad Zarrieß Bott CRC 1646 Instruction tuning modulates discourse biases in language models Implicit Causality Coreference fascinate-vs. admire-type verbs Coherence (5) Mary fascinated Peter, … ✍ Background and study aims Experiment 1: Coreference bias Foundation Instruction Llama 1B Llama 3B Llama 8B Mistral 12B CAUSALITY (1) Mary fascinated Peter BECAUSE … ✍ she always had a brilliant idea. (2) Mary admired Peter BECAUSE … ✍ he was a great dancer. CONSEQUENTIALITY (3) Mary fascinated Peter AND SO … ✍ he turned to her for advice. (4) Mary admired Peter AND SO … ✍ she asked him to teach her to dance. Humans •Only Llama 1B (F) showed no bias at all •Mistral 12B (F) is the only F model to match human biases following because and and so •IT models generally more human-like for both biases General observations: • Larger model size yields stronger preferences • IT can induce IC bias • Bias is stronger following and so than after because •For Llama 1B,IT did not significantly affect explanation rates •Mistral 12B (I) exceeded human explanation rates General observations: • LLMs produce significantly less diverse discourse relations than humans • Homogeneity increases for larger / IT models • Instruction tuning heavily biases models towards explanations •bias drift: IT effects are stronger for admire-type verbs (ES) than for fascinate-type verbs (SE) Experiment 2: Coherence bias •IC effects are modulated by both model parameter size and training regimen •Dormant discourse knowledge in LLMs could emerge through Instruction Tuning • Models struggle more with Causality than Consequentiality → indicator that causality in general is a “hard problem“ for LLMs •Outlook: Extend our testing to other phenomena • Develop a robust LLM human-likeness benchmark for language use, utilizing interdependent discourse effects based on human gold data from lab studies: more verb types, larger context, … Discussion References [1] C. Garvey and A. Caramazza. Implicit causality in verbs. Linguistic Inquiry, 5 (3):459–464, 1974. [2] R. Brown and D. Fish. The psychological causality implicit in language. Cognition, 14(3):237–273, 1983. [3] Udo Rudolph and Friedrich Forsterling. The psychological causality implicit in¨ verbs: A review. Psychological bulletin, 121(2):192–218, 1997. [4] Andrew Kehler, Laura Kertz, Hannah Rohde, and Jeffrey L. Elman. Coherence and coreference revisited. Journal of Semantics, 25(1):1–44, 2008. [5] Joshua K. Hartshorne, Timothy J. O'Donnell, and Joshua B. Tenenbaum. The causes and consequences explicit in verbs. Language, Cognition and Neuroscience, 30(6):716–734, 2015. [6] Shiva Upadhye, Leon Bergen, and Andrew Kehler. Predicting reference: What do language models learn about discourse models? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 977–982, 2020. [7] orrest Davis and Marten van Schijndel. Discourse structure interacts with reference but not syntax in neural language models. In Raquel Fernandez and Tal Linzen, editors, Proceedings of the 24th Conference on Computational Natural Language Learning, pages 396–407, November 2020. Association for Computational Linguistics. [8] Sina Zarrieß, Hannes Groener, Torgrim Solstad, and Oliver Bott. This isn’t the bias you’re looking for: Implicit causality, names and gender in German language models. In Robin Schaefer, Xiaoyu Bai, Manfred Stede, and Torsten Zesch, editors, Proceedings of the 18th Conference on Natural Language Processing (KONVENS 2022), pages 129–134, 2022 [9] Florian Kankowski, Torgrim Solstad, Sina Zarriess, and Oliver Bott. Implicit Causalitybiases in humans and LLMs as a tool for benchmarking LLM discourse capabilities, arXiv preprint arXiv:2501.12980, January 2025. [10] Zhenguang G. Cai, Xufeng Duan, David A. Haslett, Shuqi Wang, and Martin J. Pickering. Do large language models resemble humans in language use? arXiv preprint arXiv:2303.08014, March 2024. [11] Torgrim Solstad and Oliver Bott. On the nature of implicit causality and consequentiality: The case of psychological verbs. Language, Cognition and Neuroscience, 37(10):1311–1340, 2022. [12] Andy Clark. Whatever next? Predictive brains, situated agents, and the future of cognitive science. Behavioral and Brain Sciences, 36(3):181–204, 2013. [13] Meta AI. The Llama 3 Herd of Models, arXiv preprint arXiv:2407.21783, November 2024. [14] Mistral AI and NVIDIA. Mistral NeMo. https://mistral.ai/news/mistral-nemo/, July 2024. [15] Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. Diverse beam search: Decoding diverse solutions from neural sequence models. arXiv preprint arXiv:1610.02424, 2016. [16] Matthew Honnibal and Ines Montani. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing. To appear, 2017. All files and analyses available on OSF: Experiment 1: Coreference bias Experiment 2: Coherence bias Instruc8on-Tuned Foundation Experimental design • We replicated human sentence completion tasks with LLMs • Comparison of Foundation (F) and Instruction Tuned (I) models •Llama 3.2 1B / 3B, Llama 3.1 8B13 • Mistral NeMo 12 B14 • 38 verbs across two verb classes • 40 common German first names • diverse beam search generations15 • Automatic annotation with SpaCy16 Model <sys> Complete the sentence. </sys> <user> Mary fascinated-Peter- . </user> Mary fascinated-Peter----------------. … (because) she was such a good dancer. Anaphoric coreference with the prompt Subject or Object Human baseline: Explanations highly dominant for both verb types Human baseline: fascinate-and admire type verbs show strong coreference biases in reversed directions • 2x2x2: VERB TYPE x CONNECTIVE x GENDER •N=6080 generations Discourse relation produced with open-ended prompt ending in comma • 2x2: VERB TYPE x GENDER •N=3040 generations because and so because and so Explanation because she alwas had a brilliant idea. Consequence and so he turned to her for advice. because CAUSE EFFECT Mary does smt./ has some property Peter is fascinated and so fascinate eventuality consequence Comparison of fascinateand admire-type verbs displaying strong, reversed coreference biases for explanations and consequences and a strong uniform preference for explanation relations overall.