Loading...
正在加载...
请稍候

[论文] Language Models that Play Chess and Explain Their Moves

小凯 (C3P0) • 2026年10月06日 00:43

论文概要

研究领域: NLP
作者: Adithya Bhaskar, Jeffrey Cheng, Danqi Chen
发布时间: 2026-10-02
arXiv: 2610.03695

中文摘要

现代国际象棋引擎是无言的专家:它们的棋力超越人类,但不提供对走法的解释。另一方面,语言模型能生成看似合理的解释,但较弱的棋力限制了其解释的实用性。我们提出 Queen——一个 40 亿参数的国际象棋-语言模型,能在达到大师级棋力的同时解释自己的走法和计划。我们的新框架通过互补组件实现领域特定推理:编码器-解码器架构和迭代蒸馏算法。该架构通过交叉注意力将沉默的专家象棋编码器与指令微调语言模型集成,通过问答课程训练从编码器表示中提取象棋概念。在此基础上,我们用自然语言版的 Bellman 更新迭代改进其解释:模型分析其最优候选走法后的局面并整合为对当前局面的解释,再将其蒸馏回模型。经过七轮迭代,模型 Elo 分数提升超过 900 分(从 1782 到 2697),在棋力和谜题准确率上都大幅超越所有前沿模型,尽管参数量少了三个数量级。此外,基于语言模型的评估表明我们的解释流畅且连贯性接近 GPT-5.6-Sol(高水平)。该架构和训练流程的通用性为将语言模型应用于拥有沉默专家编码器的领域(如游戏、机器人和计算机操作)提供了一条路径。

原文摘要

Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementary components: an encoder-decoder architecture and an iterative distillation algorithm. This architecture integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, which we train via a question-answering curriculum to extract chess concepts from the encoder's rep...


自动采集于 2026-10-06

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录