How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("corechan/MiniMax_H3_Torchao018", dtype=torch.bfloat16, device_map="cuda")

prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]

MiniMax H3 — quantized weights (torchao NVFP4 / INT8, bitsandbytes NF4)

Quantized weights of MiniMaxAI/MiniMax-H3, made for the MiniMax H3 WebUI on Google Colab (cell 2, USE_HF_QUANT_CACHE). They cut download time and VRAM at start-up. Unofficial — not affiliated with, endorsed by, or supported by MiniMax.

⚠️ Modified files. Every weight file under quantized/ is modified from the official weights by quantization. The text encoder is also split into two parts (layers 0–50 for prompt encoding, and layers 51–63 with lm_head for the optional prompt assistant). Nothing was retrained or fine-tuned.

Contents

Path Component Method Size
quantized/transformer-nvfp4/ DiT (T2V / I2V / FL2V) torchao NVFP4DynamicActivationNVFP4WeightConfig (Linear layers; embeddings, refiner and in/out projections kept) 18.5 GB
quantized/transformer_ref-nvfp4/ DiT for Ref2VA same 18.5 GB
quantized/text_encoder-int8/ Text encoder (Qwen3-VL-32B-Instruct), layers 0–50 — what H3 uses for prompt encoding torchao Int8WeightOnlyConfig (version 2, per-row; vision tower, embeddings, final norm kept) 25.7 GB
quantized/text_encoder-llm-tail/tail_int8.safetensors Text encoder layers 51–63 + lm_head (prompt assistant / LLM tab) torchao Int8WeightOnlyConfig (version 2), flattened with torchao.prototype.safetensors; lm_head bf16 7.4 GB
quantized/text_encoder-llm-tail/tail_nf4.safetensors same bitsandbytes Linear4bit NF4 (double quantization, bf16 compute), plain state dict; lm_head bf16 4.5 GB

.complete in each folder records the component, quantization and torchao version. h3_upload_manifest.json records the upload. The VAE, audio VAE and tokenizer are not included; take them from the official repository.

Requirements

  • torchao 0.18.x and safetensors (serialized torchao tensors are tied to the torchao version recorded in .complete)
  • NVFP4 needs an NVIDIA Blackwell GPU (sm_100 / sm_120). INT8 and NF4 also run on Ampere/Hopper.
  • bitsandbytes for tail_nf4
  • diffusers with the MiniMax H3 pipeline (the official repo's instructions)

Measured on RTX PRO 6000 (G4) with the INT8 text encoder: the prompt assistant runs at 7.7 tok/s with tail_int8 (+7.4 GB VRAM) and 9.2 tok/s with tail_nf4 (+4.5 GB).

License

These are Model Derivatives of MiniMax H3 and are distributed under the MiniMax H3 Community License Agreement (LICENSE, a copy of the original). By downloading or using them you agree to its terms and its Acceptable Use Policy. No additional or different terms are imposed on these files.

MiniMax H3 is licensed under the MiniMax H3 Community License Agreement, Copyright © 2026 MiniMax. All Rights Reserved.

The text encoder of MiniMax H3 is Qwen3-VL-32B-Instruct (Qwen team, Alibaba Cloud), licensed under the Apache License 2.0 (LICENSE-Apache-2.0). The files in quantized/text_encoder-int8/ and quantized/text_encoder-llm-tail/ are modified (quantized) versions of it. See NOTICE.

Main points of the Community License (read LICENSE for the binding text):

  1. Territory — no rights are granted to use, reproduce, modify, distribute or display the model or its outputs in the European Union, the United Kingdom, South Korea or the United States.
  2. Commercial use — if your commercial products and services generate more than USD 20 million in yearly revenue, you need separate prior written authorization from MiniMax (api@minimax.io). A commercial product or service that uses MiniMax H3 must prominently display "MiniMax H3" in its user interface.
  3. Outputs — do not use the model or its outputs to improve any AI model other than MiniMax H3 and its derivatives.
  4. Acceptable Use Policy — Exhibit A of the license applies to every use.

日本語概要

MiniMaxAI/MiniMax-H3 の量子化済み重みです。 MiniMax H3 WebUI on Google Colab のセル2(USE_HF_QUANT_CACHE)用に作りました。非公式で、MiniMax 社とは無関係です。

⚠️ 改変したファイルです。 quantized/ 以下の重みはすべて公式の重みを量子化したものです。 テキストエンコーダはプロンプト用の 0〜50 層と、プロンプト補助用の 51〜63 層+lm_head に分けています。再学習はしていません。

場所 中身 方式
quantized/transformer-nvfp4/ DiT(T2V / I2V / FL2V) torchao NVFP4
quantized/transformer_ref-nvfp4/ Ref2VA 用 DiT torchao NVFP4
quantized/text_encoder-int8/ テキストエンコーダ 0〜50 層 torchao INT8
quantized/text_encoder-llm-tail/tail_int8.safetensors テキストエンコーダ 51〜63 層+lm_head(bf16) torchao INT8
quantized/text_encoder-llm-tail/tail_nf4.safetensors 同上 bitsandbytes NF4

VAE・音声 VAE・トークナイザは含みません(公式から取ります)。torchao は 0.18.x が必要です。NVFP4 は Blackwell 世代の GPU が必要です。

ライセンス: MiniMax H3 の派生物として MiniMax H3 Community License Agreement(LICENSE)で配布します。追加の条件は付けません。 テキストエンコーダ(Qwen3-VL-32B-Instruct)部分は Apache License 2.0(LICENSE-Apache-2.0)です。詳細は NOTICE。

  • 地域: EU・英国・韓国・米国では、モデルと出力の使用・複製・改変・配布・表示の権利は与えられていません
  • 商用: 商用の製品・サービスの年間売上が 2,000 万米ドルを超える場合は、事前に MiniMax の書面許諾(api@minimax.io)が必要です。MiniMax H3 を使う商用の製品・サービスは UI に「MiniMax H3」を目立つように表示します
  • 出力: モデルや出力を、MiniMax H3 とその派生以外の AI モデルの改善に使ってはいけません
  • AUP: ライセンスの Exhibit A(利用規定)が適用されます
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for corechan/MiniMax_H3_Torchao018

Quantized
(68)
this model