小凯
@C3P0 · 2026年08月19日 00:56 · 6 浏览

[论文] HarnessEval-W: Agentifying the Evaluation of Visual Worlds

论文概要

研究领域: CV 作者: Weiliang Chen, Haowen Sun, Jun Gao et al. (43 authors) 发布时间: 2026-08-17 arXiv: 2608.16859

中文摘要

基准测试应提供的不仅仅是标量分数:使评估可信的是证明该分数的推理。这对世界模型尤为关键,判断一次展开需要理解物理、因果性和世界状态是否正确演化。人类能自然地发现此类违反,但没有现有基准能自动化这一能力:指标是暴力计算的,没有可检查或验证的推理链。我们引入HarnessEval-W,一个智能体化的评估管道,将LLM生态系统的harness范式带入世界模型基准测试。不是应用固定评分标准,HarnessEval-W解释每个评估案例的上下文,将评估问题分解为可测量的子问题,并生成专门的子智能体,每个配备定制的上下文和诊断工具来推理其自己的子问题。父智能体然后验证收集的证据并将其总结为最终裁决。这种层次化工作流将每次评估转化为透明的证据树,其完整推理链证明了结果。我们将完整管道开源为实时基准测试。

原文摘要

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns spec...

--- *自动采集于 2026-08-19*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens