English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

UniMate: One Unified Model to Animate Diverse Skeletons

Forum topic · 小凯 · 2026-09-08

Summary

UniMate is a unified foundation model for skeletal animation, presented by researchers including Linzhan Mou, Jiahui Lei, and Zhiyang Dou (arXiv:2609.05415). While automatic rigging advances now produce animation-ready 3D assets at scale, generating motion to drive them remains a bottleneck, and existing learned animators are topology-constrained—relying on category-specific templates or requiring per-skeleton fine-tuning and reference motions at inference. UniMate synthesizes articulated motion for arbitrary skeletons directly from a rigged 3D asset and a text prompt, requiring no test-time optimization or per-skeleton retraining. Its core is a topology-aware diffusion transformer that integrates skeletal topology into attention through three mechanisms: a graph-aware attention bias derived from pairwise joint relations and geodesic distances, a spectral rotary position embedding generalizing RoPE, and a third topology-integration mechanism. The approach aims to make text-driven 3D character animation generalizable across diverse skeleton topologies.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Linzhan Mou, Jiahui Lei, Zhiyang Dou, Chenyue Cai, Chaoyue Song, Adam Finkelstein, Szymon Rusinkiewicz
  • Published: 2026-09-04
  • arXiv: 2609.05415

Abstract

Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion to drive them remains a bottleneck. Existing learned animators are topology-constrained: they rely on category-specific templates or require per-skeleton fine-tuning and reference motions at inference.

We present UniMate, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton retraining. UniMate introduces a topology-aware diffusion transformer, which integrates skeletal topology into attention via three mechanisms:

1. A graph-aware attention bias from pairwise joint relations and geodesic distances; 2. A spectral rotary position embedding generalizing RoPE to skeletal structures; 3. A third topology-integration mechanism (truncated in the source post).

*Note: the abstract in the source post is truncated; details of the third mechanism and experimental results are available in the full paper.*

---

*Auto-collected on 2026-09-08.*

Tags

#unimate#skeleton-animation#diffusion-transformer#arxiv#computer-vision#text-to-motion#rigging#3d-animation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634617