Abstract
UniMate is a unified diffusion transformer that generates articulated motion for arbitrary skeletons from text and rigged 3D assets without per-skeleton retraining, using topology-aware attention and a large curated motion dataset.
Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion to drive them remains a bottleneck. Existing learned animators are topology-constrained: they rely on category-specific templates or require per-skeleton fine-tuning and reference motions at inference. We present UniMate, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton retraining. UniMate introduces a topology-aware diffusion transformer, which integrates skeletal topology into attention via three mechanisms: (1) a graph-aware attention bias from pairwise joint relations and geodesic distances; (2) a spectral rotary position embedding generalizing RoPE to arbitrary kinematic trees via the graph Laplacian; and (3) a global topological conditioner attention-pooled from the rest-pose skeleton. We also curate UniML3D, 13,006 motion sequences spanning bipedal, quadrupedal, avian, marine, insectoid, serpentine, and articulated rigid objects with unified canonicalization and text pairing. Trained on this dataset, UniMate outperforms state-of-the-art baselines in quality, generalization, and efficiency, and supports zero-shot cross-topology transfer, in-betweening, expansion, and text-guided editing. Our project page is available at https://linzhanmou.com/unimate/.
Community
UniMate is a unified diffusion transformer that animates arbitrary skeletons from a rigged 3D asset and a text prompt, with no per-skeleton retraining or test-time optimization. It uses a graph-aware attention bias, a spectral RoPE built from the graph Laplacian, and a global topological conditioner pooled from the rest pose. We also release UniML3D, 13,006 canonicalized, text-paired motion sequences across bipeds, quadrupeds, birds, marine animals, insects, snakes and articulated rigid objects.
Project page: https://linzhanmou.com/unimate/
Code: https://github.com/Friedrich-M/UniMate
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Neural Motion Blending Across Arbitrary Character Topologies (2026)
- BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives (2026)
- ViP-Rig: Visual-Prompted Controllable Rigging (2026)
- 4DStreamCtrl: Interactive Video Generation with Online 4D Control (2026)
- Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation (2026)
- Kirin: Animal Motion Generation from In-the-Wild Video (2026)
- HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.05415 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 5
Linzhan/Mixamo-Animations-Characters
Spaces citing this paper 0
No Space linking this paper