Overview
Understanding human activities and their surrounding environments typically relies on visual perception, yet cameras pose persistent challenges in privacy, safety, energy efficiency, and scalability. This paper explores an alternative: 4D perception without vision, reconstructing human motion and 3D scene layouts purely from everyday wearable sensors.
- Field: Computer Vision
- Authors: Hao-Yu Hsu, Tianhang Cheng, Jing Wen
- Published: 2026-04-23
- arXiv: 2604.21934
- IMU-to-4D is a framework that repurposes large language models for non-visual spatiotemporal understanding of human-scene dynamics.
- It requires only data from a few inertial sensors (IMUs) found in earbuds, watches, or smartphones—no cameras needed.
- The model predicts detailed 4D human motion together with coarse scene structure.
- Experiments across diverse human-scene datasets show IMU-to-4D yields more coherent and temporally stable results than state-of-the-art cascaded pipelines.
Key Points
Significance
The results indicate that wearable motion sensors alone can support rich 4D understanding of humans and their environments, offering a privacy-friendly and energy-efficient alternative to camera-based perception.
---
*Auto-collected on 2026-04-25.*