论文概要
研究领域: NLP
作者: Emily Cheng, Ryan Cotterell
发布时间: 2026-08-28
arXiv: 2608.28560
中文摘要
听者能否仅从话语形式恢复说话者的意思?我们从信息论角度回答了这个问题,对于由任何文本特征提取器(包括当代大语言模型的隐藏状态)给出的听者。将语言使用建模为意义、语境和话语的联合分布,我们推导了解码器从话语表示恢复说话者意图意义的概率上界。这些界由形式留下的关于意义的不确定性控制,它分为不可约部分和仅(语言外语境)而非话语本身可以解析的部分。因为这些量是语言的内在属性,无论产生它的表示使用了多少文本或监督,都无法超越它们;这些界无论意义空间是离散还是连续都成立。人工语言、汉语零代词解析和颜色参考上的实验为理论提供了经验证据。
原文摘要
Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language use as a joint distribution over meanings, contexts, and utterances, we derive upper bounds on the probability that a decoder recovers a speaker's intended meaning from a representation of the utterance. The bounds are governed by the uncertainty that form leaves about meaning, which splits into an irreducible part and a part that only (extralinguistic) context, but never the utterance alone, can resolve. Because these quantities are intrinsic to language, no representation, however much text or supervision produced it, can s...
自动采集于 2026-09-01
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。