English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

A Formal Limitation on Learning Human Language From Textual Corpora (Cheng & Cotterell, arXiv 2608.28560)

Forum topic · 小凯 · 2026-09-01

Summary

This paper by Emily Cheng and Ryan Cotterell (arXiv:2608.28560, NLP) asks whether a listener can recover a speaker's intended meaning from the form of an utterance alone. The authors answer information-theoretically for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language use as a joint distribution over meanings, contexts, and utterances, they derive upper bounds on the probability that a decoder recovers the speaker's intended meaning from an utterance representation. The bounds are governed by the uncertainty that form leaves about meaning, which splits into an irreducible part and a part that only extralinguistic context—never the utterance alone—can resolve. Because these quantities are intrinsic properties of a language, no representation, regardless of how much text or supervision produced it, can surpass them. The bounds hold whether the meaning space is discrete or continuous. Experiments on artificial languages, Chinese zero-pronoun resolution, and color reference provide empirical evidence for the theory.

Overview

Field: NLP Authors: Emily Cheng, Ryan Cotterell arXiv: 2608.28560

Abstract

Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language use as a joint distribution over meanings, contexts, and utterances, we derive upper bounds on the probability that a decoder recovers a speaker's intended meaning from a representation of the utterance. The bounds are governed by the uncertainty that form leaves about meaning, which splits into an irreducible part and a part that only (extralinguistic) context, but never the utterance alone, can resolve. Because these quantities are intrinsic to language, no representation, however much text or supervision produced it, can surpass them; the bounds hold whether the meaning space is discrete or continuous. Experiments on artificial languages, Chinese zero-pronoun resolution, and color reference provide empirical evidence for the theory.

Key points

  • Information-theoretic upper bounds on meaning recovery from utterance form, applicable to any text featurizer including LLM hidden states.
  • Uncertainty about meaning decomposes into an irreducible component and a component resolvable only via extralinguistic context.
  • The bounds are intrinsic to the language: no amount of text or supervision in the representation can exceed them.
  • Bounds hold for both discrete and continuous meaning spaces.
  • Empirical validation on artificial languages, Chinese zero-pronoun resolution, and color reference tasks.

Tags

#nlp#arxiv#information-theory#large-language-models#semantics#meaning-recovery#theoretical-nlp

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634338