English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery for Autonomous Driving

Forum topic · 小凯 · 2026-04-23

Summary

SpanVLA is an end-to-end autonomous driving framework that integrates autoregressive vision-language reasoning with a flow-matching action expert. To address the high latency and limited robustness of existing Vision-Language-Action (VLA) models that generate actions autoregressively, SpanVLA introduces an efficient bridge that leverages VLM vision and reasoning guidance to plan future trajectories via a flow-matching policy conditioned on historical trajectory initialization, significantly reducing inference time. The authors further propose a GRPO-based post-training method that teaches the model to learn from positive samples while also avoiding typical negative behaviors and learning recovery behaviors. A new dataset, mReasoning, is introduced, focusing on complex reasoning-heavy driving scenarios with negative-recovery samples. Extensive experiments on NAVSIM v1 and v2 demonstrate competitive planning performance and robustness across diverse scenarios. Paper on arXiv: 2604.19710.

Paper Overview

Field: Computer Vision / Autonomous Driving Authors: Zewei Zhou, Ruining Yang, Xuewei Qi, Yiluan Guo, Sherry X. Chen, Tao Feng, Kateryna Pistunova, Yishan Shen, Lili Su, Jiaqi Ma arXiv: 2604.19710

Abstract

Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially in long-tail scenarios. However, existing VLA models often struggle with the high latency in action generation using an autoregressive generation framework and exhibit limited robustness.

This paper proposes SpanVLA, a novel end-to-end autonomous driving framework integrating autoregressive reasoning with a flow-matching action expert:

  • Efficient action bridging: SpanVLA introduces an efficient bridge to leverage the vision and reasoning guidance of the VLM, planning future trajectories with a flow-matching policy conditioned on historical trajectory initialization, which significantly reduces inference time.
  • Learning from negative-recovery: A GRPO-based post-training method improves performance and robustness, enabling the VLA model to learn not only from positive samples but also how to avoid typical negative behaviors and learn recovery behaviors.
  • mReasoning dataset: A new real-world driving reasoning dataset focused on complex, reasoning-intensive scenarios with negative-recovery samples.
  • Extensive experiments on NAVSIM (v1 and v2) demonstrate the competitiveness of SpanVLA, and qualitative results across diverse scenarios highlight the model's planning performance and robustness.

    Links

  • arXiv: https://arxiv.org/abs/2604.19710
*Auto-collected on 2026-04-23*

Tags

#autonomous-driving#vla#flow-matching#grpo#reinforcement-learning#papers#arxiv#computer-vision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618661