Paper Overview
- Field: Computer Vision (CV)
- Authors: Ziyi Wang, Xinshun Wang, Shuang Chen, Yang Cong, Mengyuan Liu
- Published: 2026-03-23
- arXiv: 2603.22282
Introduction
The authors present UniMotion, to their knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single architecture.
Limitations of Existing Models
Existing unified models handle only restricted modality subsets (e.g., Motion-Text or static Pose-Image) and predominantly rely on discrete tokenization, which introduces quantization errors and disrupts temporal continuity.
Method
UniMotion overcomes both limitations through a core principle: treating motion as a first-class continuous modality on equal footing with RGB.
1. Cross-Modal Aligned Motion VAE (CMA-VAE) — together with symmetric dual-path embedders, it constructs parallel continuous pathways for Motion and RGB within a shared LLM backbone. 2. Dual Posterior Alignment (DPA) — injects visual-semantic priors into motion representations without requiring images at inference time, by distilling the richer posteriors of a vision-fused encoder into the motion-only encoder. 3. Latent Reconstruction Alignment (LRA) — a self-supervised pretraining strategy addressing the cold-start problem (text supervision alone being too sparse to calibrate the newly introduced motion pathway). It uses dense motion latent representations as explicit conditions to jointly calibrate the embedders, backbone, and flow head, establishing a stable motion-aware foundation for all downstream tasks.
Results
UniMotion achieves state-of-the-art performance across seven tasks covering arbitrary-to-any understanding, generation, and editing among the three modalities.
---
*Auto-collected on 2026-03-25.*