English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving

Forum topic · 小凯 · 2026-05-23

Summary

Sensor2Sensor (arXiv 2505.17379) is a generative framework that converts monocular in-the-wild dashcam video into high-fidelity multimodal autonomous driving sensor suites, including multi-view camera images and LiDAR point clouds. The core challenge is the absence of paired training data between unstructured web video and structured AV logs. The authors address this by using 4D Gaussian Splatting (4DGS) reconstruction and novel view synthesis to render real AV logs into dashcam-style videos, creating pseudo-paired data for training a diffusion-based conversion model. Quantitative evaluations demonstrate the fidelity and realism of the generated sensor data, and the authors show practical utility in transforming internet and dashcam footage into realistic multimodal formats suitable for autonomous driving development.

Paper Overview

Field: Computer Vision (CV) Authors: Jiahao Wang, Bo Sun, Yijing Bai Published: 2025-05-23 arXiv: 2505.17379

Abstract

Robust training and validation of autonomous driving systems (ADS) requires large-scale, diverse datasets. While proprietary data collected by autonomous vehicle (AV) fleets offers high fidelity, it is limited in scale, sensor configuration diversity, and geographic and long-tail behavioral coverage. In contrast, in-the-wild data from sources such as dashcams provides massive scale and diversity, capturing critical long-tail scenarios and novel environments. However, this unstructured field video data is incompatible with the structured multimodal sensor input that ADS expect for validation and training.

To bridge this data gap, the authors propose Sensor2Sensor, a novel generative modeling paradigm that converts in-the-wild monocular dashcam video into high-fidelity multimodal sensor suites (AV logs), including multi-view camera images and LiDAR point clouds.

Key Idea

The central challenge is the lack of paired training data between field videos and AV sensor logs. The solution:

  • Use 4D Gaussian Splatting (4DGS) reconstruction and novel view rendering to convert real AV logs into dashcam-style videos.
  • Train a diffusion-based architecture on this pseudo-paired data to perform the generative conversion from dashcam video to full multimodal sensor suites.
  • Results

  • Comprehensive quantitative evaluation of the fidelity and realism of generated sensor data.
  • Practical demonstration: converting challenging in-the-wild internet and dashcam footage into realistic multimodal data formats, unlocking vast external data sources for autonomous driving development.
---

*Auto-collected on 2026-05-23.*

Tags

#autonomous-driving#computer-vision#generative-models#diffusion-models#gaussian-splatting#lidar#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620662