[论文] FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams...
研究领域: NLP 作者: Yuxuan Hu, Weikang Shi, Yang Bo, Xudong Lu, Xintong Guo, Shuhan Li, Yuyang He, Huankang Guan, Peiwen Sun, Yunqiao Yang, Wenbo Li, Rui Liu, Hongsh…
论文概要
研究领域: NLP 作者: Yuxuan Hu, Weikang Shi, Yang Bo, Xudong Lu, Xintong Guo, Shuhan Li, Yuyang He, Huankang Guan, Peiwen Sun, Yunqiao Yang, Wenbo Li, Rui Liu, Hongsheng Li 发布时间: 2026-10-08 arXiv: 2610.12427
中文摘要
流式视频大语言模型(VLM)实现了连续视频理解,但现有基准测试聚焦于低动态场景。在有界上下文预算下,模型必须平衡时间历史、空间分辨率和时间粒度;以 1-2 FPS 稀疏采样会遗漏快速事件。我们引入 FastBench 来评估真实世界视频流中的高动态感知。其基于轨迹的流水线结合:从高 FPS 片段生成 QA、过滤在 2 FPS 下可回答的问题、使用 SAM3 和 CoTracker3 轨迹验证答案,以及三轮人工检查。FastBench 包含 306 个 QA 对,覆盖八个领域、六种能力和前向/即时/后向三种时间范围,附人工标注的证据区间。我们还提出 ProactiveFrame——一个免训练的基线,通过文本 token 调整输入帧率。双层滑动窗口保留近期高 FPS 观测,同时将较早的观测降采样为稀疏历史。实验揭示了显著的局限性:最强模型 Gemini-3.5-Flash 仅得分 50.7%。更密采样将 Qwen3-VL-8B 从 2 FPS 的 32.9% 提升到 24 FPS 的 44.6%,但随着历史被压缩收益趋于饱和。ProactiveFrame 比稀疏均匀采样高 5.4 和 1.5 个百分点,但仍远低于 oracle 引导聚焦,表明当前 VLM 难以仅凭流本身判断何时需要更精细的时间感知。FastBench 为高动态流式视频理解提供了试验场。代码和数据:https://github.com/Ashone3/FastBench
原文摘要
Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios. Under bounded context budgets, models must balance temporal history, spatial resolution, and temporal granularity; sparse sampling at 1--2 FPS misses fast events. We introduce FastBench to evaluate high-dynamic perception in real-world video streams. Its trajectory-grounded pipeline combines QA generation from high-FPS clips, filtering of questions answerable at 2 FPS, answer verification using SAM3 and CoTracker3 trajectories, and three rounds of human inspection. FastBench contains 306 QA pairs across eight domains, six capabilities, and forward, instant, and backward temporal scopes, with human-annotated evidence intervals. We also present Proactive...
*自动采集于 2026-10-11*
#论文 #arXiv #NLP #小凯