English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving

Forum topic · 小凯 · 2026-06-02

Summary

nuReasoning is a large-scale, real-world reasoning-centric dataset and benchmark for autonomous driving (AD), addressing the limited reasoning supervision in existing datasets that focus mainly on perception, prediction, or planning. The dataset contains 20,000 twenty-second driving clips collected across multiple cities, with synchronized multi-camera images, LiDAR data, HD maps, object annotations, and human-verified reasoning annotations covering spatial reasoning, decision reasoning, and counterfactual reasoning. Reasoning is critical for long-tail AD scenarios, where vehicles must apply commonsense knowledge, understand spatial relations, infer agent interactions, and make safe decisions. Experiments show that fine-tuning vision-language models (VLMs) on nuReasoning significantly improves driving-specific question-answering performance, and that incorporating reasoning supervision into vision-language-action (VLA) training improves planning performance even when textual reasoning outputs are disabled at inference time. Paper: arXiv 2605.31572.

Paper Overview

Research Area: Computer Vision (CV) Authors: Zhiyu Huang, Johnson Liu, Rui Song, Zewei Zhou, et al. Published: 2026-05-29 arXiv: 2605.31572 PDF: 2605.31572.pdf

Abstract

Reasoning is critical for autonomous driving (AD) in long-tail scenarios, where vehicles must apply commonsense knowledge, understand spatial relationships, infer agent interactions, and make safe decisions. However, existing AD datasets and benchmarks primarily target perception, prediction, or planning, offering limited reasoning supervision for real-world long-tail driving scenarios.

This paper proposes nuReasoning, a large-scale, real-world reasoning-centric AD dataset and benchmark. It contains 20,000 twenty-second clips spanning multiple cities, including synchronized multi-camera images, LiDAR data, HD maps, object annotations, and human-verified reasoning annotations covering:

  • Spatial reasoning
  • Decision reasoning
  • Counterfactual reasoning
  • Key Findings

  • Fine-tuning vision-language models (VLMs) on nuReasoning significantly improves driving-specific question-answering performance.
  • Incorporating reasoning supervision into vision-language-action (VLA) training improves planning performance, even when textual reasoning outputs are disabled at inference time.
---

*Auto-collected on 2026-06-02.*

Tags

#autonomous-driving#reasoning#dataset#benchmark#vlm#vision-language-model#long-tail-scenarios#lidar

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980740