Skip to main navigation Skip to search Skip to main content

What do language models know when they know language? A theoretical framework for explaining the linguistic behavior of LLMs

  • Céline Evianne Budding

Research output: ThesisPhd Thesis 1 (Research TU/e / Graduation TU/e)

176 Downloads (Pure)

Abstract

The impressive performance of LLMs -such as their ability to produce realistic linguistic outputs and their ability to learn new tasks from a small number of examples- has led to uncertainty about how the performance of these systems can be described and whether they can and should be attributed linguistic and cognitive capacities like understanding and intelligence. What makes it difficult to assess whether these attributions are warranted is that LLMs are opaque. While we might have access to their inputs and outputs, we do not know what LLMs learn from their training data, nor how or why they produce certain outputs. This opacity can be resolved by explaining the behavior of LLMs, by showing how or why these systems return particular outputs. Yet, as the behavior of LLMs can be explained in different ways, it is difficult to compare different explanations and to draw conclusions about what LLMs can actually do. What is missing, therefore, is a theoretically-grounded explanatory framework that provides a shared vocabulary and agreed-upon terms for developing explanations of the behavior of LLMs. To this end, the central research question guiding this dissertation is: How can we explain the behavior of large language models? To answer this question, the dissertation puts forward a theoretical framework within which explanations of the linguistic behavior of LLMs can be developed. Specifically, the contributions of this dissertation are twofold. First, it motivates a cognitivist approach towards explaining the behavior of LLMs and demonstrates how Rosa Cao’s account of representational pragmatism provides clear criteria for positing genuinely explanatory representations in the context of LLMs. Second, it provides a more specific framework for explaining the linguistic behavior of LLMs through the identification of tacit knowledge in these systems. In particular, the dissertation demonstrates how Martin Davies’ notion of tacit knowledge can be adapted for and applied to contemporary LLMs. To show how the framework can be applied in practice, the dissertation concludes by evaluating two explanation methods that have been developed in the empirical literature and demonstrating that there is convincing, albeit preliminary, evidence that at least one of these methods actually identifies tacit knowledge in contemporary LLMs. The dissertation consists of seven chapters. Chapter 1 provides a general introduction to the research topic and justifies why it is important to explain the behavior of LLMs. Chapter 2 considers how linguistic behavior has been explained more generally and motivates a cognitivist approach, which posits cognitive states and processes such as mental representations. Specifically, the chapter argues that, in the context of AI, Rosa Cao’s account of representational pragmatism provides clear conditions for when a pattern of activity can be considered a genuinely explanatory representation. Chapter 3 then provides the technical background for the systems to be explained, specifically focusing on the history, architecture, and relevance of transformer-based LLMs. The tacit knowledge framework is proposed and developed in chapter 4. Tacit knowledge, in Davies’ sense, refers to knowledge that is not represented explicitly but that nevertheless plays a causally-relevant role in behavior. In this way, it provides a way to posit a kind of knowledge even in the absence of explicitly represented rules. More precisely, tacit knowledge specifies a particular causal-explanatory structure for the internal processing of a system, which explains how that system can recognize and process particular semantic patterns in the data. What makes Davies’ account of tacit knowledge particularly suitable in the context of LLMs is that it provides clear constraints for when tacit knowledge can be attributed. Taken together, chapter 4 demonstrates how Davies’ constraints for tacit knowledge can be adapted for and applied to contemporary LLMs, taking into account their specific architectural features. Chapters 5 and 6 show how the tacit knowledge framework can be used to develop explanations in practice. Specifically, chapter 5 addresses how surgical interventions can be used to identify causal common factors in a target system by showing how the output of the system depends on particular variables in its internal causal structure. Building on Woodward’s interventionist account of causal explanation, chapter 5 sets up clear criteria for evaluating whether a system exhibits the causal-explanatory structure required for being attributed tacit knowledge. To determine whether there is evidence that LLMs in fact acquire tacit knowledge, chapter 6 evaluates the results of two intervention-based explanation methods developed in the recent empirical literature. This analysis demonstrates that there is promising- albeit preliminary- evidence that some LLMs acquire tacit knowledge. Finally, chapter 7 synthesizes the results of the research presented in this dissertation and discusses the broader implications of the tacit knowledge framework, for example in the context of mitigating bias in LLMs. With regard to discussions about the capacities of LLMs, this dissertation shows that it is not only important to explain the behavior of LLMs, but also to consider how such explanations can best be developed in the first place. The tacit knowledge framework proposed in this dissertation contributes to this last question by providing clear terms and criteria for developing cognitivist explanations of the behavior of LLMs. In doing so, it sets a first step towards better explaining the behavior of LLMs and towards eventually resolving debates about their potential linguistic and/or cognitive capacities.
Original languageEnglish
QualificationDoctor of Philosophy
Awarding Institution
  • Industrial Engineering and Innovation Sciences
Supervisors/Advisors
  • Müller, Vincent C., Promotor
  • Zednik, Carlos, Copromotor
Award date26 Nov 2025
Place of PublicationEindhoven
Publisher
Print ISBNs978-94-6522-695-8
Publication statusPublished - 26 Nov 2025

Bibliographical note

Proefschrift.

Fingerprint

Dive into the research topics of 'What do language models know when they know language? A theoretical framework for explaining the linguistic behavior of LLMs'. Together they form a unique fingerprint.

Cite this