论文概要
研究领域: NLP
作者: Honghao Lin, David P. Woodruff, Yuan Deng, Jieming Mao, Song Zuo, Vahab Mirrokni
发布时间: 2026-09-14
arXiv: 2609.15983
中文摘要
语言模型可以生成看似合理的短证明,但在长视野研究问题上仍可能不可靠——在这类问题中,进展依赖于一系列不确定且相互关联的决策。我们提出 Stellar Colosseum,一个模型无关的推理分配框架,用于数学和理论计算机科学领域的研究。Colosseum 在构建证明之前探索替代策略,使用"就绪门控"来判断一条路径是否成熟到可以分解,将证明计划表示为相互关联的章节级子问题,并将验证器的发现路由回论证中受影响的部分。在这些阶段中,系统并行生成候选方案,用针对性的证伪进行攻击,并通过重叠随机样本树聚合将候选方案及其批评合并为单一研究产物。Colosseum 工作流已集成到 Google Antigravity 的 Teamwork 框架中,作为"长证明模式"。我们通过开放式研究和定理证明与竞赛编程基准评估展示了 Colosseum 的能力。使用 Gemini 3.1 Pro,我们在 FOCS 和 JMLR 等顶会论文提出的开放问题上取得了若干新结果。在 TCS-Bench(来自 FOCS、STOC 和 SODA 论文的研究级定理证明基准)上,Colosseum 使用 Gemini 3.1 Pro 和 Gemini 3.7 Flash 达到 71.0% 的准确率。在另一项使用 Gemini 3.1 Pro 的 Codeforces 评估中,带执行反馈的面向证明的流水线解决了 222 题中的 218 题。
原文摘要
Language models can produce plausible short proofs, but may still be unreliable on long-horizon research problems, where progress depends on a sequence of uncertain and interdependent decisions. We introduce Stellar Colosseum, a model-agnostic harness for allocating inference across research in mathematics and theoretical computer science. Colosseum explores alternative strategies before proof construction, uses a readiness gate to decide when a route is mature enough to decompose, represents the proof plan as interdependent section-level subproblems, and routes verifier findings back to the affected part of the argument. Across these stages, it generates candidates in parallel, attacks them with targeted falsification, and combines candidates and their critiques into a single research art...
自动采集于 2026-09-16
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。