Overview
Field: NLP Authors: Emily Cheng, Ryan Cotterell arXiv: 2608.28560
Abstract
Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language use as a joint distribution over meanings, contexts, and utterances, we derive upper bounds on the probability that a decoder recovers a speaker's intended meaning from a representation of the utterance. The bounds are governed by the uncertainty that form leaves about meaning, which splits into an irreducible part and a part that only (extralinguistic) context, but never the utterance alone, can resolve. Because these quantities are intrinsic to language, no representation, however much text or supervision produced it, can surpass them; the bounds hold whether the meaning space is discrete or continuous. Experiments on artificial languages, Chinese zero-pronoun resolution, and color reference provide empirical evidence for the theory.
Key points
- Information-theoretic upper bounds on meaning recovery from utterance form, applicable to any text featurizer including LLM hidden states.
- Uncertainty about meaning decomposes into an irreducible component and a component resolvable only via extralinguistic context.
- The bounds are intrinsic to the language: no amount of text or supervision in the representation can exceed them.
- Bounds hold for both discrete and continuous meaning spaces.
- Empirical validation on artificial languages, Chinese zero-pronoun resolution, and color reference tasks.