Text-to-Video
Diffusers
MiniMax H3
MiniMaxH3Pipeline
custom_minimax_h3
video-generation
image-to-video
audio-video-generation
quantized
4-bit precision
int4
svdquant
quantfunc
comfyui
Instructions to use QuantFunc/Minimax-H3-Quantfunc-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use QuantFunc/Minimax-H3-Quantfunc-4bit with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("QuantFunc/Minimax-H3-Quantfunc-4bit", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Add RELEASE-NOTES-2026-09-19.md
Browse files- RELEASE-NOTES-2026-09-19.md +17 -0
RELEASE-NOTES-2026-09-19.md
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# MiniMax-H3 QuantFunc 4-bit — release 2026-09-19 (v11)
|
| 2 |
+
|
| 3 |
+
Both files in this repo were replaced on 2026-09-19. Anyone who downloaded earlier files will NOT see the change — this note is the record.
|
| 4 |
+
|
| 5 |
+
| file | build | replaces |
|
| 6 |
+
|---|---|---|
|
| 7 |
+
| `minimax_h3_fl2va_4step_quantfunc_int4_r128.safetensors` | fl2va, svdq int4 r128 **g64**, lightx2v **v1.1** 4-step turbo LoRA merged into every slot **including the token_refiner**, **int4 token_refiner** (`refiner_a4w4_v1`), **int8-conv lowrank sidecar** (`sidecar_int8_v1`, per-slot `int8_conv`) | same name; old file 12,297,885,392 B sha256 c665d9c60eb1210de1bef91e9ef50d395b879a4ee9d79bc45c57b5ba7ab9ea30 (pre-g64 export) |
|
| 8 |
+
| `minimax_h3_ref2va_8steps_quantfunc_int4_r128.safetensors` | ref2va, svdq int4 r128 **g64**, lightx2v **v1.0 8-step** turbo LoRA merged incl. token_refiner, **int4 token_refiner**, **int8-conv lowrank sidecar** | NEW name (8 steps); the old `minimax_h3_ref2va_4step_quantfunc_int4_r128.safetensors` (13,143,213,304 B, sha256 b425621a20b1e2e797dbf50c5fe48fe0cd3deb28f902a6f4197ff574bfdc4663, pre-g64) is SUPERSEDED — do not use it; it is scheduled for removal by the repo owner |
|
| 9 |
+
|
| 10 |
+
What changed vs the previous files: (1) g64 int4 law; (2) the fp16 token_refiner is now LoRA-merged — every earlier export ran the HF
|
| 11 |
+
base refiner unmerged (exporter defect, fixed); (3) the token_refiner is quantized to int4 as well (measured indistinguishable from the
|
| 12 |
+
fp16-refiner build over 6 seeds on both models; less VRAM); (4) fl2va uses the v1.1 LoRA (v1.0 before).
|
| 13 |
+
|
| 14 |
+
Use: ComfyUI-QuantFunc native loader; ref2va = 8 sampling steps, fl2va = 4. Per-slot precision is read from the sealed metadata (no switches).
|
| 15 |
+
Acceptance (bf16+LoRA baseline, same box/session): fl2va first-frame fine-detail retention 0.93 of bf16+v1.1 (3 seeds); on the v1.0 line
|
| 16 |
+
over 6 seeds 0.923 of bf16 (int8 convrot 0.988, previous 4-bit public file class 0.83); ref2va locked-composition wall-scene detail 0.97–0.98
|
| 17 |
+
of bf16. Full evidence lives in the QuantFunc fleet dossier `issue-qfc-h3-ref2ab-lightx2v-export`.
|