Overview
Field: Computer Vision (CV) Authors: Junke Wang, Qihang Zhang, Shuai Yang Published: 2025-06-13 arXiv: 2506.10666
Abstract
This work presents RepWAM, a representation-centric world action model (WAM) built on representation visual-action tokenizers. Existing WAMs typically inherit reconstruction-oriented video tokenizers from pretrained video generation models. Although these tokenizers preserve visual fidelity, pixel reconstruction alone provides limited guidance for learning instruction-following dynamics that connect future prediction with robot control. To address this, the authors explore a semantic visual-action latent space for representation-centric world action modeling.
Key points
- Train a representation visual-action tokenizer that maps visual inputs into aligned visual and latent action tokens.
- Pretrain the WAM to jointly model future visual states and the latent actions that connect them, under language instructions.
- Adapt the pretrained model to real robot trajectories for closed-loop manipulation.
- Experiments on real-world manipulation tasks and simulation benchmarks show strong performance across various manipulation settings.
- Ablation studies highlight the value of semantic visual-action tokenization over reconstruction-oriented alternatives.
- The results establish representation visual-action tokenization as a promising foundation for world action models, moving toward generalist robot policies.
*Automatically collected on 2026-06-14.*