English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Evidence-Backed Video Question Answering: Grounding Video LLM Answers with Spatio-Temporal Evidence

Forum topic · 小凯 · 2026-07-15

Summary

A new paper on arXiv (2607.11862) proposes Evidence-Backed Video Question Answering (E-VQA), a task requiring Video Large Language Models to output both a semantic answer and verifiable spatio-temporal evidence: temporal segments plus dense, tracked object segmentation masklets. Existing explainability approaches rely on textual rationales or sparse bounding boxes that fail to capture complex dynamics like occlusions and non-rigid deformations. The authors introduce ST-Evidence, the first human-verified benchmark covering both discriminative and generative pixel-level grounding, and find that QA accuracy in state-of-the-art models is critically decoupled from true visual perception—a gap that scaling alone cannot close. They also build an automated pipeline to create the 160k-sample ST-Evidence-Instruct dataset; fine-tuned grounding Video LLMs substantially outperform the same-sized UniPixel baseline (t-mean +27.2, J&F +13.8 on 7B models), establishing a robust baseline for explainable, evidence-based video understanding.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Shijie Wang, Honglu Zhou, Ziyang Wang, Ran Xu, Caiming Xiong
  • Published: 2026-07-13
  • arXiv: 2607.11862
  • Summary

    Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations.

    The paper proposes Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise spatio-temporal evidence: temporal segments and dense, tracked object segmentation masklets.

    Key Contributions

  • ST-Evidence benchmark: the first human-verified benchmark for both discriminative and generative pixel-level grounding.
  • Key finding: evaluations of state-of-the-art models reveal a critical decoupling between QA accuracy and genuine visual perception—one that cannot be closed by scaling alone.
  • ST-Evidence-Instruct dataset: a scalable automated generation pipeline producing 160k training samples.
  • Strong results: fine-tuning grounding Video LLMs on this dataset yields large improvements over the same-sized UniPixel baseline (t-mean +27.2, J&F +13.8 on 7B models), establishing a robust baseline for explainable, evidence-based video understanding.
---

*Auto-collected on 2026-07-15.*

Tags

#video-llm#question-answering#visual-grounding#explainability#computer-vision#segmentation#benchmark#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178395147