论文概要
研究领域: CV
作者: Yu Chen, Caorui Li, Ziyu Xiong
发布时间: 2026-07-22
arXiv: 2507.17085
中文摘要
长音频-视频推理对全模态大模型而言具有挑战性,因为决定性证据通常稀疏、跨模态,且用统一高保真输入保留成本过高。我们引入\textbf{OmniReasoner},一个用于长音频-视频思考的工具使用后训练框架:全模态大模型通过监督微调和强化学习,学会在回答前决定是否以及何时调用放大工具。OmniReasoner首先构建完整流的低成本全局预览,然后在需要时调用放大工具,请求特定时间间隔进行更高保真度的视觉和音频检查后再回答。由于模型在调用前后观察到不同的采样粒度——稀疏的全局预览和更稠密的局部片段——我们引入\textbf{TimeAnchor},使工具的时间参数在不同粒度间保持有效和往返一致,而非绑定到特定采样率的帧索引。为使这种工具使用行为可训练而无需昂贵的手动区间标注,我们构建时间增强数据引擎,通过视频编辑和合成为工具使用后训练合成轨迹。跨全模态和视频基准的实验表明,OmniReasoner在将高保真计算集中于信息丰富区域的同时,提高了回答准确率和时间定位能力。代码已开源。
原文摘要
Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve with uniformly high-fidelity inputs. We introduce OmniReasoner, a tool-use post-training framework for Thinking with Long Audio-Video: omni-modal LLMs learn, via supervised fine-tuning and reinforcement learning, to decide whether and where to call a zoom-in tool before answering. OmniReasoner first builds a low-cost global preview of the full stream and then, when needed, calls the zoom-in tool with a requested temporal interval for higher-fidelity visual and audio inspection before answering. Because the model observes different sampling granularities before and after this call -- a sparse global preview and a denser local clip -- we introdu...
自动采集于 2026-07-23
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。