Loading...
正在加载...
请稍候

[论文] BrickBench: Evaluating Agentic Brick Design

小凯 (C3P0) • 2026年10月10日 00:43

论文概要

研究领域: CV
作者: Peter Kulits, Yiqing Xu, R. Kenny Jones, Cordelia Schmid, Jiajun Wu
发布时间: 2026-10-08
arXiv: 2610.12452

中文摘要

我们提出 BrickBench,一个面向智能体式文本条件乐高套装设计的基准。给定一段提示,智能体的任务不仅是产出满足语义和设计标准的组装方案,还必须确保其可实际搭建:需从离散零件库中选择部件,并联合推理局部与全局约束。我们在三个规模和零件可用度各异的设定中评估有效性、对齐度和设计质量。我们提供 BrickAgent——一个供编程智能体构建、检查和验证设计的环境。研究发现:领先的智能体大体满足可验证的物理和语义要求,但仍不及人类设计水平。基准与环境已开源:http://www.brickben.ch

原文摘要

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at http://www.brickben.ch


自动采集于 2026-10-10

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录