VerMo Model Card

A Motius-native autoregressive motion-language baseline aligned on HumanML3D captions.

Motius Implementation | Motius Checkpoint

VerMo is a Motius-native research baseline rather than a reproduction of an external paper. The released M2T checkpoint uses a Llama-3.2-1B-Instruct language backbone, a 16K motion tokenizer, and an explicit SMPL-22 motion processor. No external paper or original repository is claimed for this row.

Release Snapshot

Item Value
Released task M2T
Evaluation input HumanML3D-263, converted through the Motius SMPL-22 bridge
Native motion representation VerMo-138, 20 fps
Motion tokenizer 16K VQ motion tokenizer
Language backbone Llama-3.2-1B-Instruct
Checkpoint ZeyuLing/Motius-VerMo-HumanML3D
Pipeline motius.pipelines.vermo.VermoPipeline

Usage

import numpy as np
from motius.pipelines.vermo import VermoPipeline

pipe = VermoPipeline.from_pretrained(
    "ZeyuLing/Motius-VerMo-HumanML3D",
    bundle_kwargs={"device": "cuda"},
    smpl_model_dir="checkpoints/body_models/smpl",
)
motion = np.load("sample.npy")  # denormalized HumanML3D-263
caption = pipe.infer_m2t([motion], lengths=[len(motion)])[0]

M2T Evaluation

Protocol Samples BLEU-4 ROUGE-L CIDEr BERT F1 R@1 R@2 R@3 Matching
HumanML3D M2T 4,400 - - - - - - - -

Motion Representation

VerMo-138 stores absolute root translation (3), frame-to-frame root translation (3), and 22 local joint rotations in column-major 6D form (132). HumanML3D inputs are recovered to SMPL-22 joints, solved to motion135 with position IK, then repacked explicitly from row-major to column-major 6D.

Motius Components

Component Path
Pipeline motius/pipelines/vermo/pipeline.py
Bundle motius/models/vermo/bundle.py
Processor motius/models/vermo/processor.py
Motion tasks motius/models/vermo/task_utils/
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support