Summary
MoTok is a diffusion-based discrete motion tokenizer for human motion proposed by Chenyang Gu, Mingyuan Zhang, and Haozhe Xie (arXiv:2503.16903). The key idea is to decouple semantic abstraction from fine-grained motion reconstruction: the tokenizer learns discrete tokens that capture high-level semantics, while a diffusion decoder handles the recovery of detailed motion. This separation allows the representation to bridge semantic conditions (e.g., text descriptions) and kinematic conditions (e.g., trajectories) within a single framework. On the HumanML3D benchmark, MoTok significantly improves controllability and fidelity while using only one-sixth of the tokens required by prior approaches, reducing trajectory error from 0.72 cm to 0.08 cm. The work falls in the computer vision field and targets motion generation/editing tasks where both semantic control and precise kinematic detail matter.
Overview
Field: Computer Vision (CV)
Authors: Chenyang Gu, Mingyuan Zhang, Haozhe Xie
Published: 2026-03-19
arXiv: 2503.16903
Abstract
We propose MoTok, a diffusion-based discrete motion tokenizer that decouples semantic abstraction from fine-grained reconstruction by delegating motion recovery to a diffusion decoder. On HumanML3D, our method significantly improves controllability and fidelity while using only one-sixth of the tokens, reducing trajectory error from 0.72 cm to 0.08 cm.
Key contributions
- A diffusion-based discrete motion tokenizer that separates semantic abstraction from fine-grained motion reconstruction.
- A diffusion decoder responsible for recovering detailed motion from the discrete tokens.
- A unified representation bridging semantic conditions (e.g., text) and kinematic conditions (e.g., trajectories).
Results
- Benchmark: HumanML3D
- Uses only 1/6 of the tokens compared to prior methods
- Trajectory error reduced from 0.72 cm to 0.08 cm
*Auto-collected on 2026-03-22.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177168983