English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving

Forum topic · 小凯 · 2026-06-12

Summary

VLGA is a vision-language-action (VLA) model for autonomous driving that grounds driving actions in a dense 3D world by adding geometry as a fourth modality alongside vision, language, and action. A dedicated geometry expert is supervised with a per-pixel pointmap regression loss against LiDAR, ensuring the policy actively uses dense spatial information rather than relying on frozen 3D foundation features or sparse box/map losses. On open-loop nuScenes, VLGA achieves state-of-the-art among ego-state-free VLA methods with an average L2 of 0.50 m and a 3-second collision rate of 0.18%. In closed-loop evaluation on Bench2Drive, it reaches a driving score of 79.08, exceeding the previous strongest VLA by 0.71 with comparable efficiency and comfort. The paper (arXiv 2606.12396) is authored by Jin Yao, Dhruva Dixith Kurra, Tom Lampo, Zezhou Cheng, Danhua Guo, and Burhan Yaman.

Paper Overview

  • Field: Computer Vision (autonomous driving)
  • Authors: Jin Yao, Dhruva Dixith Kurra, Tom Lampo, Zezhou Cheng, Danhua Guo, Burhan Yaman
  • Published: 2026-06-10
  • arXiv: 2606.12396
  • Abstract

    Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing approaches either inject features from a frozen 3D foundation model without an objective that ensures the policy uses them, or constrain geometry with sparse box and map losses that provide no dense spatial signal.

    Key Contributions

  • VLGA is introduced as the first vision-language-action model supervised to reconstruct the dense 3D world it drives through.
  • Geometry is added as a fourth modality alongside vision, language, and action, handled by a dedicated expert supervised by a per-pixel pointmap regression loss against LiDAR.
  • Results

    Extensive experiments were conducted on the challenging nuScenes and Bench2Drive datasets for open-loop and closed-loop evaluation respectively, showing VLGA's superiority over corresponding VLA baselines:

  • Open-loop (nuScenes): New state of the art among ego-state-free VLA methods, with the lowest L2 error (average 0.50 m) and the lowest 3-second collision rate (0.18%).
  • Closed-loop (Bench2Drive): State-of-the-art driving score of 79.08, +0.71 over the previously strongest VLA, with comparable efficiency and comfort.
---

*Auto-collected on 2026-06-12.*

Tags

#vlga#autonomous-driving#vision-language-action#3d-geometry#pointmap#nuscenes#bench2drive#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981121