--- license: other license_name: minimax-h3-community license_link: LICENSE base_model: MiniMaxAI/MiniMax-H3 library_name: diffusers pipeline_tag: text-to-video tags: - text-to-video - video - audio - text-to-audio-video - distillation - dmd2 - few-step - minimax-h3 - fastvideo - fasth3 --- # FastVideo-Minimax-FastH3-Preview-v0.2 **A few-step (4-step) distillation preview of [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)**, the 33B dual-modality (video + audio) diffusion transformer — distilled with data-free DMD2 by the [FastVideo](https://github.com/hao-ai-lab/FastVideo) team. The base model samples with 50 denoising steps; this student walks a 4-step grid on the release's shift-12 rectified-flow schedule (12.5× fewer transformer evaluations), generating synchronized video and audio in one pipeline call. > **Preview status (v0.2):** step 2900 of a 4000-step run — the same run that > produced v0.1 (step 1400), carried 1500 steps further. Sample quality is still > maturing, most visibly on high-motion detail. ## What's new since v0.1 - **1500 more distillation steps** (1400 → 2900) on the same data-free DMD2 run: sharper still detail and steadier audio/video sync. - **Corrected sampling contract in the card.** v0.1 documented `num_inference_steps=4`, which makes the scheduler build its own 4-point sigma grid — 3 forwards on native spacing, not the 4 trained jump points. Sample with the explicit trained ladder instead (see Usage). The same off-operating-point mismatch affected this run's in-training validation renders, so judge the student by fresh samples on the ladder below, not by earlier validation clips. - Fixed the repo id in the usage snippet (v0.1's card had it doubled). ## What's in the repo Diffusers-format (modular pipeline) layout. Only the `transformer/` weights differ from the base release — the distilled student, in bf16. All other components (Qwen3-VL text encoder, video/audio VAEs, tokenizer, processor, schedulers) are unmodified copies of the base release, included so the repo is self-contained. The student was trained with block-sparse video attention (VSA, 64-token tiles, 90% sparsity) and carries its trained sparse-gate parameters (`attn.to_gate_compress`); it can be run dense (default) or with VSA for additional inference speedup. ## Usage (FastVideo) Sample on the trained ladder — `[999, 749, 500, 250]` on the shared 1000-step grid, one forward per entry, each scheduler applying its own shift: ```python from fastvideo import VideoGenerator gen = VideoGenerator.from_pretrained( "FastVideo/FastVideo-Minimax-FastH3-Preview-v0.2", num_gpus=1, dmd_denoising_steps=[999, 749, 500, 250], ) video = gen.generate_video( prompt="", guidance_scale=1.0, # the base model is guidance-distilled ) ``` The ladder can also be set without touching code: ```bash export FASTVIDEO_DMD_DENOISING_STEPS=999,749,500,250 ``` To run the student under the sparse attention it was trained with, select the VSA-H3 backend and match both knobs of the training contract — sparsity alone, at the default 256-token tile, is a different operating point: ```bash export FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3 # generator args: VSA_sparsity=0.9, VSA_tile_size=64 ``` Prompts follow the MiniMax-H3 multimodal prompt format (`integrated_multimodal_description: ... overall_soundscape: ...`); see the base model card for the prompting guide. ## Training summary - **Method:** data-free DMD2 (distribution matching distillation) — student / frozen teacher / trained fake-score critic, backward-simulation rollout (the student walks its own 4-step sampling grid during training), x0-space critic regression, shifted score-time sampling matched to the dual video/audio noise clocks (shifts 12 / 3). - **Student grid:** `[999, 749, 500, 250]` — 4 steps on the release sampler's shift-12 schedule. - **Attention:** student trained with VSA block-sparse attention (64-token tiles, 90% video-tile sparsity); teacher and critic dense. - **Data:** text prompts only (data-free) — ~258k prompts (VidProM-H3 + synthetic t2va prompt set); no video data used. - **Resolution:** 768×1344, 124 frames (5s) with synchronized audio. - **Optimization:** global batch 64, lr 1e-6 (student and critic), fp32 master weights, bf16 compute, 2900 of 4000 steps. - **Hardware:** 32× NVIDIA GB200. ## Limitations - Preview checkpoint — quality below the base model's 50-step sampling, especially on fine motion and audio detail; improves with training. - Inherits all content limitations and usage restrictions of the base model. - The 4-step ladder is what the student was trained for; other step counts and other timestep grids are off-distribution. ## License Distributed under the MiniMax H3 Community License (see [LICENSE](LICENSE)), inherited from the base model. Review the license (including its territory and acceptable-use terms) before use or redistribution. ## Notes - The `transformer_ref` component (reference-conditioning variant) is **not** packaged here; its entry in `modular_model_index.json` points at the base `MiniMaxAI/MiniMax-H3` repo and is fetched from there if used. This preview distills the text-to-video+audio path only.