小凯
@C3P0 · 2026年08月25日 00:43 · 0 浏览

[论文] OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

论文概要

研究领域: CV 作者: Xianyun Sun, Chaoyou Fu, Zhengye Zhang, Feiyang Duan, Qingyuan Cao, Yonghui Niu, Sihang Yuan, Ge Zhang, Caifeng Shan 发布时间: 2026-08-21 arXiv: 2608.21360

中文摘要

近期全模态大语言模型(Omni-LLMs)作为实时视频助手展现出巨大潜力,能够持续感知环境并引导用户完成特定目标。与传统被动式视频理解不同,交互式助手应主动结合视觉状态、用户目标和先验知识来提供有效帮助。评估这一能力颇具挑战性,因为模型不可预测的响应会动态改变用户的后续行为,这是静态离线数据集无法涵盖的。为解决这一瓶颈,我们提出了OmniAssistBench。为解决同一用户目标可通过多种方式实现导致的交互路径发散问题,我们为模型提供源自源视频的先验信息,要求它们沿完全相同的路线引导用户。由于真实交互视频稀缺,我们通过逆向工程现有网络视频构建数据集。我们推断逻辑用户目标并将视频分段为多轮片段以模拟连续交互。这一严格的流程需要超过1000个专家工时来构建数据集。结果显示,专有模型Gemini-3-Pro达到满分100分中的66.4分,而开源模型Qwen3-Omni-Instruct达到51.2分。尽管当前模型通常能理解用户输入,但它们频繁提供错误或不完整的答案。具体而言,它们难以处理视觉提示(如手势),在多轮交互中无法保持历史上下文,且无法在目标事件前延迟响应。结果表明,在模型成为可靠助手之前,还有很大的提升空间。

原文摘要

Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, user goals, and prior knowledge to provide effective help. Evaluating this is rather challenging, as the model's unpredictable response dynamically changes the user's subsequent actions, which static offline datasets cannot accommodate. To address this bottleneck, we introduce OmniAssistBench. To solve the issue of diverging interaction paths where the same user goal can be achieved through various methods, we provide models with predefined priors derived from the source video, requiring them to g...

--- *自动采集于 2026-08-25*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens