English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OmniAssistBench: A Benchmark for Assistant-Style Interaction in Omni-LLMs

Forum topic · 小凯 · 2026-08-25

Summary

OmniAssistBench is a new benchmark for evaluating omni-modal large language models (Omni-LLMs) as real-time interactive video assistants. Unlike passive video understanding, an assistant must combine visual states, user goals, and prior knowledge to actively guide users toward goals. Because a model's responses dynamically change user behavior, static offline datasets cannot capture this interaction. The benchmark addresses diverging interaction paths by providing models with priors derived from source videos and requiring them to guide users along an identical route. Since real interactive videos are scarce, the dataset was built by reverse-engineering existing web videos: inferring logical user goals and segmenting videos into multi-turn clips, requiring over 1,000 expert hours. Results show proprietary Gemini-3-Pro scores 66.4 out of 100, while open-source Qwen3-Omni-Instruct scores 51.2. Models often understand input but give wrong or incomplete answers, struggle with visual cues like gestures, fail to maintain multi-turn context, and cannot delay responses until goal-relevant events occur.

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

  • Field: Computer Vision (CV)
  • Authors: Xianyun Sun, Chaoyou Fu, Zhengye Zhang, Feiyang Duan, Qingyuan Cao, Yonghui Niu, Sihang Yuan, Ge Zhang, Caifeng Shan
  • Released: 2026-08-21
  • arXiv: 2608.21360
  • Abstract

    Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, user goals, and prior knowledge to provide effective help.

    Evaluating this capability is challenging, because a model's unpredictable responses dynamically change the user's subsequent actions — something static offline datasets cannot accommodate. To address this bottleneck, the authors introduce OmniAssistBench.

    To solve the issue of diverging interaction paths (where the same user goal can be achieved through various methods), models are provided with predefined priors derived from the source video, requiring them to guide the user along exactly the same route.

    Because real interactive videos are scarce, the dataset was constructed by reverse-engineering existing web videos: inferring logical user goals and segmenting videos into multi-turn clips to simulate continuous interaction. This rigorous pipeline required over 1,000 expert hours.

    Results

  • Gemini-3-Pro (proprietary): 66.4 / 100
  • Qwen3-Omni-Instruct (open source): 51.2 / 100
  • Although current models generally understand user input, they frequently provide wrong or incomplete answers. Specifically, they:

  • struggle with visual cues such as hand gestures;
  • fail to maintain historical context across multi-turn interactions;
  • cannot delay responses until before a goal-relevant event.
The results indicate substantial room for improvement before models can become reliable real-time assistants.

---

*Auto-collected on 2026-08-25.*

Tags

#omni-llm#benchmark#video-assistant#multimodal#computer-vision#interactive-evaluation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633966