English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VLK: Learning Humanoid Loco-Manipulation from Synthetic Vision-Language-Kinematics Interactions

Forum topic · 小凯 · 2026-07-01

Summary

This paper introduces VLK, a framework for perception-based humanoid loco-manipulation that maps egocentric observations and language instructions to whole-body motion. Since no existing dataset provides synchronized egocentric images, language commands, and robot-compatible kinematic trajectories at scale, the authors generate vision-language-kinematics (VLK) supervision synthetically in reconstructed scenes. Their pipeline uses 3D Gaussian Splatting to reconstruct metric-scale indoor environments, synthesizes navigation and object-interaction trajectories from privileged scene information, and renders paired egocentric observations afterward. Without any human intervention, they produce 48,000 paired trajectories and train a VLK policy that predicts short-horizon whole-body kinematic trajectories, which a whole-body tracker converts into actions on a physical humanoid. Evaluated on a physical Unitree G1 performing navigation and single-object transport tasks, the results show that synthetic interactions in reconstructed scenes provide effective supervision for sim-to-real, perception-based humanoid loco-manipulation. Source: arXiv:2507.00001.

Paper Overview

Field: Robotics Authors: Yen-Jen Wang, Jiaman Li, Sirui Chen Published: 2026-07-01 arXiv: 2507.00001

Abstract

Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapping requires synchronized egocentric images, language commands, and robot-compatible kinematic trajectories, yet no existing data source provides this complete tuple at scale.

The authors address this bottleneck by generating vision-language-kinematics (VLK) supervision synthetically in reconstructed scenes. The pipeline:

1. Leverages 3D Gaussian Splatting to reconstruct metric-scale indoor environments. 2. Synthesizes navigation and object-interaction trajectories using privileged scene information. 3. Renders paired egocentric observations after the fact.

They produce 48,000 paired trajectories with no human intervention and train a VLK policy that predicts short-horizon whole-body kinematic trajectories. A whole-body tracker then converts these predictions into actions on a physical humanoid.

Results

The system was evaluated on a physical Unitree G1 executing navigation and single-object transport tasks, demonstrating that synthetic interactions in reconstructed scenes provide effective supervision for sim-to-real, perception-based humanoid loco-manipulation.

Original Abstract (excerpt)

> Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapping requires synchronized egocentric images, language commands, and robot-compatible kinematic trajectories, yet no existing data source provides this complete tuple at scale. We address this bottleneck by generating vision-language-kinematics (VLK) supervision synthetically in reconstructed scenes. Our pipeline leverages 3D Gaussian Splatting to reconstruct metric-scale indoor environments, synthesizes navigation and object-interaction trajectories using privileged scene information, and renders paired egocentric observations after the fact. We produce 48,000 paired trajectories with no human intervention and train a VLK policy that predicts...

---

*Auto-collected on 2026-07-01. Paper link: https://arxiv.org/abs/2507.00001*

Tags

#robotics#humanoid-robots#loco-manipulation#3d-gaussian-splatting#sim-to-real#machine-learning#arxiv#unitree-g1

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208336