English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation

Forum topic · 小凯 · 2026-03-25

Summary

UniMotion (arXiv:2603.22282) is presented as the first unified framework capable of simultaneously understanding and generating human motion, natural language, and RGB images within a single architecture. Unlike prior unified models that handle only limited modality subsets (e.g., Motion-Text or static Pose-Image) and rely on discrete tokenization—which causes quantization errors and disrupts temporal continuity—UniMotion treats motion as a first-class continuous modality on equal footing with RGB. It introduces a Cross-Modal Aligned Motion VAE (CMA-VAE) with symmetric dual-path embedders that build parallel continuous pathways for motion and RGB inside a shared LLM backbone. Dual Posterior Alignment (DPA) distills richer visual-semantic posteriors from a vision-fused encoder into the motion-only encoder, injecting visual priors without requiring images at inference. To address cold-start issues, Latent Reconstruction Alignment (LRA) provides self-supervised pretraining using dense motion latents to jointly calibrate embedders, backbone, and flow head. The framework achieves state-of-the-art results across seven arbitrary-to-any tasks spanning understanding, generation, and editing among the three modalities.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Ziyi Wang, Xinshun Wang, Shuang Chen, Yang Cong, Mengyuan Liu
  • Published: 2026-03-23
  • arXiv: 2603.22282

Introduction

The authors present UniMotion, to their knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single architecture.

Limitations of Existing Models

Existing unified models handle only restricted modality subsets (e.g., Motion-Text or static Pose-Image) and predominantly rely on discrete tokenization, which introduces quantization errors and disrupts temporal continuity.

Method

UniMotion overcomes both limitations through a core principle: treating motion as a first-class continuous modality on equal footing with RGB.

1. Cross-Modal Aligned Motion VAE (CMA-VAE) — together with symmetric dual-path embedders, it constructs parallel continuous pathways for Motion and RGB within a shared LLM backbone. 2. Dual Posterior Alignment (DPA) — injects visual-semantic priors into motion representations without requiring images at inference time, by distilling the richer posteriors of a vision-fused encoder into the motion-only encoder. 3. Latent Reconstruction Alignment (LRA) — a self-supervised pretraining strategy addressing the cold-start problem (text supervision alone being too sparse to calibrate the newly introduced motion pathway). It uses dense motion latent representations as explicit conditions to jointly calibrate the embedders, backbone, and flow head, establishing a stable motion-aware foundation for all downstream tasks.

Results

UniMotion achieves state-of-the-art performance across seven tasks covering arbitrary-to-any understanding, generation, and editing among the three modalities.

---

*Auto-collected on 2026-03-25.*

Tags

#paper#arxiv#computer-vision#human-motion#multimodal#unified-model#motion-generation#llm

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169027