English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Seeing Without Eyes: IMU-to-4D Reconstructs Human Motion and 3D Scenes from Wearable IMUs

Forum topic · 小凯 · 2026-04-27

Summary

A forum post introduces IMU-to-4D, a research paper (arXiv:2604.21926) by Hao-Yu Hsu, Tianhang Cheng, Jing Wen, Alexander G. Schwing, and Shenlong Wang in computer vision. The work explores camera-free 4D perception: reconstructing human motion and 3D scene layout purely from wearable inertial measurement units (IMUs) found in headsets, watches, or smartphones. IMU-to-4D repurposes large language models for non-visual spatiotemporal human-scene dynamics understanding, predicting detailed 4D human motion and coarse scene structure from a small number of IMU sensors. This avoids the privacy, security, energy-efficiency, and scalability challenges associated with camera-based perception. Experiments across diverse human-scene datasets show the framework produces more coherent and temporally stable results than state-of-the-art cascaded pipelines, demonstrating that wearable motion sensors alone can support rich 4D human-scene understanding.

Overview

This forum post shares a computer vision paper:

  • Title: Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs
  • Authors: Hao-Yu Hsu, Tianhang Cheng, Jing Wen, Alexander G. Schwing, Shenlong Wang
  • arXiv: 2604.21926
  • Key Idea

    Understanding human activities and their surroundings typically relies on visual perception, but cameras pose persistent challenges around privacy, security, energy efficiency, and scalability. The authors explore an alternative: camera-free 4D perception, aiming to reconstruct human motion and 3D scene layout purely from everyday wearable sensors.

    The IMU-to-4D Framework

    IMU-to-4D repurposes large language models for non-visual spatiotemporal human-scene dynamics understanding. It takes data from a small number of inertial sensors—such as those in headsets, watches, or smartphones—and predicts:

  • Detailed 4D human motion
  • Coarse 3D scene structure

Results

Experiments on diverse human-scene datasets show that IMU-to-4D produces more coherent and temporally stable results than state-of-the-art (SoTA) cascaded pipelines, indicating that wearable motion sensors alone can support rich 4D understanding.

---

*Auto-collected on 2026-04-27.*

Tags

#computer-vision#wearable-imu#large-language-models#4d-reconstruction#human-motion#scene-understanding#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618795