Paper Overview
Field: Robotics Authors: Yen-Jen Wang, Jiaman Li, Sirui Chen Published: 2026-07-01 arXiv: 2507.00001
Abstract
Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapping requires synchronized egocentric images, language commands, and robot-compatible kinematic trajectories, yet no existing data source provides this complete tuple at scale.
The authors address this bottleneck by generating vision-language-kinematics (VLK) supervision synthetically in reconstructed scenes. The pipeline:
1. Leverages 3D Gaussian Splatting to reconstruct metric-scale indoor environments. 2. Synthesizes navigation and object-interaction trajectories using privileged scene information. 3. Renders paired egocentric observations after the fact.
They produce 48,000 paired trajectories with no human intervention and train a VLK policy that predicts short-horizon whole-body kinematic trajectories. A whole-body tracker then converts these predictions into actions on a physical humanoid.
Results
The system was evaluated on a physical Unitree G1 executing navigation and single-object transport tasks, demonstrating that synthetic interactions in reconstructed scenes provide effective supervision for sim-to-real, perception-based humanoid loco-manipulation.
Original Abstract (excerpt)
> Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapping requires synchronized egocentric images, language commands, and robot-compatible kinematic trajectories, yet no existing data source provides this complete tuple at scale. We address this bottleneck by generating vision-language-kinematics (VLK) supervision synthetically in reconstructed scenes. Our pipeline leverages 3D Gaussian Splatting to reconstruct metric-scale indoor environments, synthesizes navigation and object-interaction trajectories using privileged scene information, and renders paired egocentric observations after the fact. We produce 48,000 paired trajectories with no human intervention and train a VLK policy that predicts...
---
*Auto-collected on 2026-07-01. Paper link: https://arxiv.org/abs/2507.00001*