--- license: apache-2.0 datasets: - FastVideo/Wan2.2-Syn-121x704x1280_32k base_model: - Wan-AI/Wan2.2-TI2V-5B-Diffusers library_name: fastvideo tags: - video-generation - text-to-video - wan - int8 - quantization - apple-silicon pipeline_tag: text-to-video --- # FastMetal-5B-QAD **3-step text-to-video, INT8 pre-quantized for Apple Silicon.** The mid-tier FastMetal model — a DMD2-distilled Wan2.2 TI2V 5B with a quantization-aware-trained INT8 DiT. 720p-native, pre-quantized: no startup quantization. ## What's inside | Path | Contents | |---|---| | `mlx_dit.safetensors` / `mlx_dit.json` | INT8 (affine, group-64) DiT | | `text_encoder/`, `vae/`, `tokenizer/`, `scheduler/` | everything needed to run standalone | ## Quickstart Requires macOS with Apple silicon (MPS) and Python 3.11+: ```bash pip install torch transformers mlx safetensors av imageio imageio-ffmpeg git clone https://github.com/FastVideo/FastVideo.git cd FastVideo python examples/inference/basic/mlx_wan22_generate.py \ --text-encoder-root ./FastMetal-5B-QAD \ --mlx-checkpoint ./FastMetal-5B-QAD \ --vae-root ./FastMetal-5B-QAD/vae \ --prompt "a river winding through a fantasy valley at golden hour" \ --fast ``` ## Model details | | | |---|---| | Base model | Wan 2.2 TI2V 5B | | Distillation | DMD2, 3 denoising steps | | Quantization | affine INT8, group size 64, QAT-trained | | Resolution | 704×1280 (720p), 121 frames | | Flow shift | 5.0 | | DiT weights | ~5 GB (INT8) | ## Training DMD2 distillation of the Wan 2.2 TI2V 5B teacher onto an INT8 student on NVIDIA GB200 clusters, with quantization-aware training (affine INT8, group 64). Training corpus: `FastVideo/Wan2.2-Syn-121x704x1280_32k`. ## FastMetal family | Model | Tier | |---|---| | [FastMetal-1.3B-QAD](https://huggingface.co/FastVideo/FastMetal-1.3B-QAD) | Entry — 16 GB+ class Macs | | **FastMetal-5B-QAD** | Mid — 720p | 16 GB+ class Macs | | [FastMetal-14B-QAD](https://huggingface.co/FastVideo/FastMetal-14B-QAD) | Quality — 24 GB+/ Ideally 36 Macs |