English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs (IMU-to-4D)

Forum topic · 小凯 · 2026-04-25

Summary

IMU-to-4D is a new framework from researchers including Hao-Yu Hsu, Tianhang Cheng, and Jing Wen that reconstructs 4D human motion and 3D scene layout without any cameras. Instead of visual perception—which raises privacy, safety, energy, and scalability concerns—the method repurposes large language models for non-visual spatiotemporal understanding of human-scene dynamics. Using only a few inertial measurement units (IMUs) commonly found in earbuds, smartwatches, or smartphones, IMU-to-4D predicts detailed 4D human motion together with coarse scene structure. Experiments across diverse human-scene datasets show that the approach produces more coherent and temporally stable results than state-of-the-art cascaded pipelines. The work, listed as arXiv 2604.21934 in the computer vision field, suggests that everyday wearable motion sensors alone can support rich 4D perception of people and their environments.

Overview

Understanding human activities and their surrounding environments typically relies on visual perception, yet cameras pose persistent challenges in privacy, safety, energy efficiency, and scalability. This paper explores an alternative: 4D perception without vision, reconstructing human motion and 3D scene layouts purely from everyday wearable sensors.

  • Field: Computer Vision
  • Authors: Hao-Yu Hsu, Tianhang Cheng, Jing Wen
  • Published: 2026-04-23
  • arXiv: 2604.21934
  • Key Points

  • IMU-to-4D is a framework that repurposes large language models for non-visual spatiotemporal understanding of human-scene dynamics.
  • It requires only data from a few inertial sensors (IMUs) found in earbuds, watches, or smartphones—no cameras needed.
  • The model predicts detailed 4D human motion together with coarse scene structure.
  • Experiments across diverse human-scene datasets show IMU-to-4D yields more coherent and temporally stable results than state-of-the-art cascaded pipelines.

Significance

The results indicate that wearable motion sensors alone can support rich 4D understanding of humans and their environments, offering a privacy-friendly and energy-efficient alternative to camera-based perception.

---

*Auto-collected on 2026-04-25.*

Tags

#arxiv#computer-vision#imu#wearables#large-language-models#human-motion-capture#3d-scene-understanding#4d-perception

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618732