[论文] FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluatin...
研究领域: ML 作者: Hoyoung Lee, Suyeol Yun, Jack Haverty, Yunju Cho, Meesong Kim, Daekyung Park, Sumin Kim, Jihoon Kwon, Jasmine Jia Geng, Andrew Chin, Yin Luo, Edwa…
论文概要
研究领域: ML 作者: Hoyoung Lee, Suyeol Yun, Jack Haverty, Yunju Cho, Meesong Kim, Daekyung Park, Sumin Kim, Jihoon Kwon, Jasmine Jia Geng, Andrew Chin, Yin Luo, Edward Tong, Yu Yu, Zach Golkhou, Minkyu Kim, Igor Halperin, Young Cha, Alejandro Lopez-Lira, Chanyeol Choi, Yongjae Lee 发布时间: 2026-09-28 arXiv: 2609.35744
中文摘要
评估金融研究智能体需要反映专家标准并固定信息截止日正确值的评分量规(rubric)。专家审核的金融基准依赖固定的逐项量规,扩展成本高且无法编码各机构的自有标准。在 FinAutoRubric 中,专家指定可复用的评估指导,而智能体和代码执行针对具体查询的量规生成、审核和验证。专家指导通过提示和代码强制执行的规则来约束每个智能体,可复用标准的任务库将其跨任务传递。在长视野循环中,编写智能体研究每个期望值,审核智能体验证它,失败时升级给人工。在三个专家编写的金融基准上,其量规在追踪专家评分方面与最强的评估生成器一样紧密,同时覆盖更多标准的专家量规期望值;其分数与人类评分一致,内部分析师在盲审中更偏好它们。发布的 100 查询 FinAutoRubric 基准来自内部分析师在 78 个任务和八个资产类别上的关键问题,表明较早模型生成的量规仍为较新模型留有改进空间。
原文摘要
Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are costly to extend and cannot encode each institution's own standard. In FinAutoRubric, experts specify reusable evaluation guidance, while agents and code carry out query-specific rubric generation, review, and validation. This expert guidance governs every agent, as prompts and as rules that code enforces, and a Task Bank of reusable criteria carries it across tasks. In long-horizon loops that follow the expert guidance, a writer agent researches every expected value and a reviewer agent verifies it, and failures escalate to a human. On three expert-authored finance benchmarks,...
*自动采集于 2026-09-30*
#论文 #arXiv #ML #小凯