English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Diagnosing Directional Motion Blindness in Video-LLMs: MoDirect Dataset and DeltaDirect Method

Forum topic · 小凯 · 2026-05-25

Summary

This paper from Jongseo Lee, Hyuntak Lee, and Sunghun Kim (arXiv:2505.14483) identifies 'directional motion blindness' in video large language models (Video-LLMs): despite advances in temporal video understanding, most models perform near chance level on simple videos where a single object moves left, right, up, or down, with apparent above-chance results largely attributable to prediction bias rather than genuine direction understanding. The authors trace motion direction information through the Video-LLM pipeline, finding that direction remains linearly accessible in the visual encoder, projector, and LLM hidden states, but the readout layer fails to bind this signal to the correct linguistic answer—a 'direction binding gap.' They introduce MoDirect, a dataset family for motion-direction instruction tuning and evaluation, and DeltaDirect, a projector-level objective that predicts normalized 2D motion vectors from adjacent-frame feature differences. On MoDirect-SynBench, DeltaDirect instruction tuning lifts motion direction accuracy from 25.9% to 85.4%; on MoDirect-RealBench it improves real-world accuracy by 21.9 percentage points over a vanilla baseline without real-world fine-tuning data, while preserving standard video understanding performance.

Paper Overview

Field: Computer Vision (CV) Authors: Jongseo Lee, Hyuntak Lee, Sunghun Kim arXiv: 2505.14483

Key Findings

  • Video-LLMs have progressed rapidly in temporal video understanding, yet many fail at a basic perceptual primitive: signed image-plane motion direction.
  • On simple videos of a single object moving left, right, up, or down, most Video-LLMs perform near random chance. Apparent above-chance performance is largely explained by prediction bias rather than true direction understanding. The authors call this failure directional motion blindness.
  • Probing the Video-LLM pipeline shows motion direction remains linearly accessible in the visual encoder, projector, and LLM hidden states, but the readout layer fails to bind this signal to the correct linguistic answer option—revealing a direction binding gap.
  • Synthetic motion-direction instruction fine-tuning can narrow the gap in-domain, but motion-direction concept vector analysis shows visual complexity attenuates the signal magnitude and limits out-of-domain generalization.
  • Proposed Solutions

  • MoDirect: a dataset family for motion-direction instruction tuning and evaluation.
  • DeltaDirect: a diagnostic-driven, projector-level objective that predicts normalized 2D motion vectors from adjacent-frame feature differences.
  • Results

  • On MoDirect-SynBench, instruction tuning with DeltaDirect improves motion-direction accuracy from 25.9% to 85.4%.
  • On MoDirect-RealBench, DeltaDirect improves real-world motion-direction accuracy by 21.9 percentage points over the vanilla baseline—without any real-world fine-tuning data—while preserving standard video understanding performance.
---

*Auto-collected on 2026-05-25*

Tags

#video-llms#computer-vision#motion-direction#instruction-tuning#multimodal-models#benchmark#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620756