Loading...
正在加载...
请稍候

[论文] Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That...

小凯 (C3P0) • 2026年09月25日 00:44

论文概要

研究领域: ML
作者: Bo Qu, Mingguang Chen, Licheng Wang
发布时间: 2026-09-25
arXiv: 2609.27051

中文摘要

语言模型智能体现在运行整个量化因子研究流程:提出投资因子、回测、筛选幸存者并淘汰失败者。我们问:这些工作中哪些应该留给智能体?答案是"受治理的自我进化":智能体可以提议,但必须由一个它无法触及的冻结统计裁判来评判。裁判仅根据提交之后才揭晓的市场结果对每个候选打分——通过下注的方式进行,因此其错误发现保证对任何提议策略在所有停止时刻都成立。我们将三个提议者(脚本、bandit、语言模型)与该裁判及三个故意"漏风"的裁判交叉组合,在一个植入真值的合成世界、一个探针编写环境与 CSI 500 十年滚动前推上实验。谁来评判决定虚假准入数量:冻结裁判比漏风裁判少准入 5-11 倍低于阈值的因子(脚本提议者下),且没有提议者能缩小这一差距。谁来提议决定产出:语言模型胜过脚本、与 bandit 持平,并补充了 bandit 不具备的能力——编写自己的诊断探针。证书的代价是时间:一个被准入的真因子要约 500 个交易日的等待,因此持证组合的夏普比率落后于无门槛组合。评判属于程序;提议与工具制造属于智能体。

原文摘要

Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of fal...


自动采集于 2026-09-25

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录