English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction with 4D Radar-Camera Fusion

Forum topic · 小凯 · 2026-07-14

Summary

4DR360 is a 4D radar-camera framework for 360-degree full-scene perception in autonomous driving, proposed by Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, and Hui-liang Shen (arXiv:2607.09629). The method addresses the sparsity of 4D millimeter-wave radar returns by fusing them with camera data, and models semantic occupancy as a persistent scene state rather than a terminal output. The framework follows a cross-modal state reasoning paradigm with two key modules: State-guided BEV Enhancement (SBE), which strengthens intra-frame BEV representations, and Doppler-guided Temporal Fusion (DTF), which preserves state evidence over longer time horizons. The authors also extend the ManTruckScenes dataset with a unified cross-dataset detection-occupancy protocol. This work targets reliable autonomous driving perception that jointly captures foreground objects and dense semantic layout.

Paper Overview

Research Area: CV / Autonomous Driving

Authors: Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen

Published: 2026-07-10

arXiv: 2607.09629

Summary

Reliable autonomous driving requires full-scene perception that couples foreground objects with a dense semantic layout. 4D millimeter-wave radar has emerged as a robust and cost-effective sensor, but its sparse returns make radar-camera fusion necessary.

This paper proposes 4DR360, a 4D radar-camera framework for 360° full-scene perception that models semantic occupancy as a persistent scene state rather than a terminal output.

The framework follows a cross-modal state reasoning paradigm, consisting of:

  • State-guided BEV Enhancement (SBE): enhances intra-frame BEV representations.
  • Doppler-guided Temporal Fusion (DTF): preserves state evidence over a longer temporal horizon.
The authors also extend the ManTruckScenes dataset with a unified cross-dataset detection-occupancy protocol.

---

*Auto-collected on 2026-07-14*

Tags

#4d-radar#autonomous-driving#3d-object-detection#occupancy-prediction#sensor-fusion#radar-camera#bev#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178395119