论文概要
研究领域: NLP
作者: Md Shohel Arman, Igor Molybog
发布时间: 2026-09-25
arXiv: 2609.31587
中文摘要
我们研究自然语言文档是否真的能帮助编程智能体解决软件问题,并构建了构造与评估此类文档的工具。我们提出一个“回环”(roundtrip)基准:用“从描述重新生成的代码能否通过原始测试”来为代码描述打分,并证明决定描述保真度的是完整性而非长度。以该基准作为优化信号,我们找到了一个能写出满保真度描述、且能泛化到未见文件的描述撰写提示词。随后我们检验了驱动这项工作的假设:更好的文档能帮助智能体解决真实仓库中的 issue。跨两个模型家族、十个仓库,并设置阳性对照以确认我们的评测确实能检测到真实改进——结果表明并不行。当源代码存在时,无论是静态精简文档还是检索到的上下文,都不优于仅提供 issue 本身。我们报告了这一负面结果,连同基准与优化器一并发布,并刻画了文档真正能够起作用的边界条件。
原文摘要
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen files. We then test the hypothesis that motivated the work: that better documentation helps an agent resolve real repository issues. Across two model families and ten repositories, and against a positive control confirming that our evaluation can detect a genuine improvement, we find that it does not. W...
自动采集于 2026-09-29
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。