--- title: Quantization Explorer emoji: ⚙️ colorFrom: blue colorTo: indigo sdk: static pinned: false license: mit short_description: "Explore quantization: FP8, INT8, INT4 and trade-offs." --- # Quantization Explorer **Quantization Explorer** is an educational Hugging Face Space by the [`open-weight`](https://huggingface.co/open-weight) organization. It explains how model quantization reduces memory requirements by representing weights at lower precision, and how common approaches such as **FP8, INT8, INT4, bitsandbytes, GPTQ, AWQ and GGUF quantization** differ in purpose and trade-offs. ## What you can explore - What model quantization is - FP16/BF16 vs. FP8 vs. INT8 vs. INT4 - Theoretical raw weight memory - Post-training quantization - On-the-fly quantization - Calibration-based methods - bitsandbytes - GPTQ - AWQ - GGUF / llama.cpp quantization - Quality, speed and compatibility trade-offs - A simple quantization decision helper ## Core idea ```text Higher-precision weights ↓ Quantization method ↓ Lower-bit representation ↓ Lower memory / storage ↓ Potential speed benefits + Possible quality / compatibility trade-offs ``` ## Primary references - Hugging Face Transformers — Quantization overview: https://huggingface.co/docs/transformers/quantization/overview - bitsandbytes: https://huggingface.co/docs/transformers/en/quantization/bitsandbytes - GPTQ: https://huggingface.co/docs/transformers/quantization/gptq - AWQ: https://huggingface.co/docs/transformers/quantization/awq - llama.cpp quantization: https://github.com/ggml-org/llama.cpp/tree/master/tools/quantize ## Related organization Open Weight https://huggingface.co/open-weight ## Related project Open Weights https://huggingface.co/open-weights ## Collaboration Open-weight AI, model infrastructure, inference, deployment, research and ecosystem partnerships. **Contact:** agenten@magenta.de