Upload 2 files
Browse files- README.md +71 -0
- index.html +366 -0
README.md
ADDED
|
@@ -0,0 +1,71 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: Quantization Explorer
|
| 3 |
+
emoji: ⚙️
|
| 4 |
+
colorFrom: blue
|
| 5 |
+
colorTo: indigo
|
| 6 |
+
sdk: static
|
| 7 |
+
pinned: false
|
| 8 |
+
license: mit
|
| 9 |
+
short_description: "Explore quantization: FP8, INT8, INT4 and trade-offs."
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
# Quantization Explorer
|
| 13 |
+
|
| 14 |
+
**Quantization Explorer** is an educational Hugging Face Space by the [`open-weight`](https://huggingface.co/open-weight) organization.
|
| 15 |
+
|
| 16 |
+
It explains how model quantization reduces memory requirements by representing weights at lower precision, and how common approaches such as **FP8, INT8, INT4, bitsandbytes, GPTQ, AWQ and GGUF quantization** differ in purpose and trade-offs.
|
| 17 |
+
|
| 18 |
+
## What you can explore
|
| 19 |
+
|
| 20 |
+
- What model quantization is
|
| 21 |
+
- FP16/BF16 vs. FP8 vs. INT8 vs. INT4
|
| 22 |
+
- Theoretical raw weight memory
|
| 23 |
+
- Post-training quantization
|
| 24 |
+
- On-the-fly quantization
|
| 25 |
+
- Calibration-based methods
|
| 26 |
+
- bitsandbytes
|
| 27 |
+
- GPTQ
|
| 28 |
+
- AWQ
|
| 29 |
+
- GGUF / llama.cpp quantization
|
| 30 |
+
- Quality, speed and compatibility trade-offs
|
| 31 |
+
- A simple quantization decision helper
|
| 32 |
+
|
| 33 |
+
## Core idea
|
| 34 |
+
|
| 35 |
+
```text
|
| 36 |
+
Higher-precision weights
|
| 37 |
+
↓
|
| 38 |
+
Quantization method
|
| 39 |
+
↓
|
| 40 |
+
Lower-bit representation
|
| 41 |
+
↓
|
| 42 |
+
Lower memory / storage
|
| 43 |
+
↓
|
| 44 |
+
Potential speed benefits
|
| 45 |
+
+
|
| 46 |
+
Possible quality / compatibility trade-offs
|
| 47 |
+
```
|
| 48 |
+
|
| 49 |
+
## Primary references
|
| 50 |
+
|
| 51 |
+
- Hugging Face Transformers — Quantization overview: https://huggingface.co/docs/transformers/quantization/overview
|
| 52 |
+
- bitsandbytes: https://huggingface.co/docs/transformers/en/quantization/bitsandbytes
|
| 53 |
+
- GPTQ: https://huggingface.co/docs/transformers/quantization/gptq
|
| 54 |
+
- AWQ: https://huggingface.co/docs/transformers/quantization/awq
|
| 55 |
+
- llama.cpp quantization: https://github.com/ggml-org/llama.cpp/tree/master/tools/quantize
|
| 56 |
+
|
| 57 |
+
## Related organization
|
| 58 |
+
|
| 59 |
+
Open Weight
|
| 60 |
+
https://huggingface.co/open-weight
|
| 61 |
+
|
| 62 |
+
## Related project
|
| 63 |
+
|
| 64 |
+
Open Weights
|
| 65 |
+
https://huggingface.co/open-weights
|
| 66 |
+
|
| 67 |
+
## Collaboration
|
| 68 |
+
|
| 69 |
+
Open-weight AI, model infrastructure, inference, deployment, research and ecosystem partnerships.
|
| 70 |
+
|
| 71 |
+
**Contact:** agenten@magenta.de
|
index.html
ADDED
|
@@ -0,0 +1,366 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
<!doctype html>
|
| 2 |
+
<html lang="en">
|
| 3 |
+
<head>
|
| 4 |
+
<meta charset="utf-8">
|
| 5 |
+
<meta name="viewport" content="width=device-width,initial-scale=1">
|
| 6 |
+
<meta name="description" content="Explore AI model quantization: FP16, BF16, FP8, INT8, INT4, bitsandbytes, GPTQ, AWQ and GGUF. Compare memory, quality and deployment trade-offs.">
|
| 7 |
+
<meta name="theme-color" content="#07111f">
|
| 8 |
+
<title>Quantization Explorer — FP8, INT8, INT4 & Open-Weight Deployment</title>
|
| 9 |
+
<style>
|
| 10 |
+
:root{
|
| 11 |
+
--bg:#07111f;--panel:#0d1b2d;--panel2:#10233a;--text:#edf7ff;--muted:#9fb4c8;
|
| 12 |
+
--line:#24445f;--cyan:#55d9ff;--blue:#6b8cff;--green:#79f2c0;--gold:#ffd580;
|
| 13 |
+
--red:#ff9e9e;--shadow:0 18px 60px rgba(0,0,0,.25)
|
| 14 |
+
}
|
| 15 |
+
*{box-sizing:border-box}
|
| 16 |
+
html{scroll-behavior:smooth}
|
| 17 |
+
body{
|
| 18 |
+
margin:0;color:var(--text);
|
| 19 |
+
font:16px/1.65 Inter,ui-sans-serif,system-ui,-apple-system,BlinkMacSystemFont,"Segoe UI",sans-serif;
|
| 20 |
+
background:
|
| 21 |
+
radial-gradient(circle at 12% 0%,rgba(85,217,255,.12),transparent 29%),
|
| 22 |
+
radial-gradient(circle at 90% 12%,rgba(107,140,255,.13),transparent 26%),
|
| 23 |
+
var(--bg)
|
| 24 |
+
}
|
| 25 |
+
a{color:var(--cyan);text-decoration:none}
|
| 26 |
+
a:hover{text-decoration:underline}
|
| 27 |
+
.wrap{max-width:1180px;margin:auto;padding:0 22px}
|
| 28 |
+
.hero{padding:72px 0 35px}
|
| 29 |
+
.badge{display:inline-flex;padding:7px 12px;border:1px solid var(--line);border-radius:999px;background:rgba(13,27,45,.75);color:#c8efff;font-size:14px}
|
| 30 |
+
h1{font-size:clamp(42px,7vw,78px);line-height:1;letter-spacing:-.055em;margin:20px 0;max-width:980px}
|
| 31 |
+
.gradient{background:linear-gradient(90deg,var(--cyan),#b6c3ff);-webkit-background-clip:text;background-clip:text;color:transparent}
|
| 32 |
+
.lead{font-size:clamp(18px,2.2vw,24px);max-width:900px;color:#cbdbe8;margin:0 0 28px}
|
| 33 |
+
.cta{display:flex;gap:12px;flex-wrap:wrap}
|
| 34 |
+
.btn{display:inline-block;padding:11px 16px;border:1px solid var(--line);border-radius:12px;font-weight:750}
|
| 35 |
+
.btn.primary{border:0;color:white;background:linear-gradient(135deg,#157aa8,#5269df)}
|
| 36 |
+
.quick{display:grid;grid-template-columns:repeat(4,1fr);gap:14px;margin:32px 0 56px}
|
| 37 |
+
.card,.section{border:1px solid var(--line);background:linear-gradient(180deg,rgba(16,35,58,.93),rgba(10,25,42,.93));border-radius:20px;box-shadow:var(--shadow)}
|
| 38 |
+
.card{padding:18px}
|
| 39 |
+
.card strong{display:block;font-size:23px;color:#fff}
|
| 40 |
+
.card span,.muted{color:var(--muted)}
|
| 41 |
+
.section{padding:28px;margin:22px 0}
|
| 42 |
+
.eyebrow{text-transform:uppercase;letter-spacing:.14em;font-size:12px;color:var(--cyan);font-weight:850}
|
| 43 |
+
h2{font-size:clamp(28px,4vw,45px);letter-spacing:-.03em;margin:6px 0 10px}
|
| 44 |
+
h3{font-size:22px;margin:0 0 8px}
|
| 45 |
+
.grid2{display:grid;grid-template-columns:1fr 1fr;gap:18px}
|
| 46 |
+
.grid3{display:grid;grid-template-columns:repeat(3,1fr);gap:14px}
|
| 47 |
+
.grid4{display:grid;grid-template-columns:repeat(4,1fr);gap:14px}
|
| 48 |
+
.callout{padding:15px 17px;border-left:4px solid var(--cyan);border-radius:9px;background:rgba(85,217,255,.06);margin:18px 0}
|
| 49 |
+
.warning{border-left-color:var(--gold);background:rgba(255,213,128,.06)}
|
| 50 |
+
.diagram{padding:22px;border:1px solid #2d5876;border-radius:17px;background:#091829;overflow:auto;margin:20px 0}
|
| 51 |
+
.flow{display:flex;gap:9px;align-items:center;min-width:900px}
|
| 52 |
+
.node{min-width:135px;padding:14px 12px;text-align:center;border:1px solid #34617d;background:#102842;border-radius:13px;font-weight:800}
|
| 53 |
+
.arrow{font-size:24px;color:var(--cyan)}
|
| 54 |
+
.tablewrap{overflow:auto}
|
| 55 |
+
table{width:100%;border-collapse:collapse;min-width:760px}
|
| 56 |
+
th,td{padding:13px;border-bottom:1px solid #23435e;text-align:left;vertical-align:top}
|
| 57 |
+
th{font-size:12px;text-transform:uppercase;letter-spacing:.07em;color:#c9efff}
|
| 58 |
+
.pill{display:inline-block;padding:5px 9px;margin:3px;border:1px solid #315b78;border-radius:999px;background:#102842;color:#c8ecff;font-size:13px}
|
| 59 |
+
.code{white-space:pre-wrap;padding:17px;border:1px solid #223f58;border-radius:14px;background:#06101c;color:#bfeeff;font:14px/1.6 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace;overflow:auto}
|
| 60 |
+
.barwrap{display:grid;grid-template-columns:repeat(5,1fr);gap:12px;align-items:end;height:220px;margin:28px 0 42px}
|
| 61 |
+
.bar{position:relative;border-radius:12px 12px 4px 4px;background:linear-gradient(180deg,var(--cyan),#5368de);min-height:28px}
|
| 62 |
+
.bar b{position:absolute;top:10px;left:0;right:0;text-align:center;color:#06111f}
|
| 63 |
+
.bar small{position:absolute;bottom:-30px;left:0;right:0;text-align:center;color:var(--muted)}
|
| 64 |
+
.controls{display:grid;grid-template-columns:1fr 1fr;gap:16px}
|
| 65 |
+
label{display:block;color:#c8efff;font-weight:700;margin-bottom:7px}
|
| 66 |
+
input,select,button{
|
| 67 |
+
width:100%;background:#0a1c2f;color:var(--text);border:1px solid #315b78;border-radius:11px;padding:11px 12px;font:inherit
|
| 68 |
+
}
|
| 69 |
+
.result{margin-top:18px;padding:18px;border:1px solid #315b78;background:#091a2b;border-radius:15px}
|
| 70 |
+
.big{font-size:34px;font-weight:850;color:white}
|
| 71 |
+
.tabs{display:flex;gap:9px;flex-wrap:wrap;margin:16px 0}
|
| 72 |
+
.tab{width:auto;cursor:pointer;border:1px solid #315b78;background:#0c2035;color:#d9f2ff;border-radius:999px;padding:8px 12px;font-weight:700}
|
| 73 |
+
.tab.active{background:linear-gradient(135deg,#177aa6,#4e66d8);border-color:transparent}
|
| 74 |
+
.answer{padding:18px;border:1px solid #315b78;border-radius:15px;background:#0a1c2f;min-height:112px}
|
| 75 |
+
.good{color:var(--green);font-weight:800}
|
| 76 |
+
.caution{color:var(--gold);font-weight:800}
|
| 77 |
+
.bad{color:var(--red);font-weight:800}
|
| 78 |
+
footer{padding:50px 0 68px;color:var(--muted)}
|
| 79 |
+
@media(max-width:820px){
|
| 80 |
+
.quick,.grid2,.grid3,.grid4,.controls{grid-template-columns:1fr}
|
| 81 |
+
.hero{padding-top:48px}.section{padding:20px}
|
| 82 |
+
.barwrap{grid-template-columns:repeat(5,minmax(52px,1fr))}
|
| 83 |
+
}
|
| 84 |
+
</style>
|
| 85 |
+
</head>
|
| 86 |
+
<body>
|
| 87 |
+
<div class="wrap">
|
| 88 |
+
<header class="hero">
|
| 89 |
+
<div class="badge">Open Weight · Quantization Explorer</div>
|
| 90 |
+
<h1>Make models smaller. Understand the <span class="gradient">trade-offs.</span></h1>
|
| 91 |
+
<p class="lead">A practical guide to model quantization — from FP16 and BF16 to FP8, INT8, INT4, bitsandbytes, GPTQ, AWQ and GGUF-based local inference.</p>
|
| 92 |
+
<div class="cta">
|
| 93 |
+
<a class="btn primary" href="#basics">Start exploring</a>
|
| 94 |
+
<a class="btn" href="https://huggingface.co/open-weight" target="_blank" rel="noopener">Open Weight organization ↗</a>
|
| 95 |
+
</div>
|
| 96 |
+
</header>
|
| 97 |
+
|
| 98 |
+
<div class="quick">
|
| 99 |
+
<div class="card"><strong>FP8</strong><span>8-bit floating-point workflows</span></div>
|
| 100 |
+
<div class="card"><strong>INT8</strong><span>Lower-memory integer inference</span></div>
|
| 101 |
+
<div class="card"><strong>INT4</strong><span>High compression for deployment</span></div>
|
| 102 |
+
<div class="card"><strong>Trade-offs</strong><span>Memory · quality · speed · support</span></div>
|
| 103 |
+
</div>
|
| 104 |
+
|
| 105 |
+
<section class="section" id="basics">
|
| 106 |
+
<div class="eyebrow">01 · Foundation</div>
|
| 107 |
+
<h2>What is model quantization?</h2>
|
| 108 |
+
<p><strong>Quantization reduces the precision used to represent model weights or activations.</strong> The goal is usually to reduce memory requirements and make models easier or cheaper to run while preserving as much model quality as possible.</p>
|
| 109 |
+
<div class="callout">Hugging Face describes quantization as lowering model memory requirements by storing weights at lower precision while trying to preserve accuracy.</div>
|
| 110 |
+
<div class="diagram">
|
| 111 |
+
<div class="flow">
|
| 112 |
+
<div class="node">FP32 / BF16 / FP16</div><div class="arrow">→</div>
|
| 113 |
+
<div class="node">Quantization method</div><div class="arrow">→</div>
|
| 114 |
+
<div class="node">FP8 / INT8 / INT4</div><div class="arrow">→</div>
|
| 115 |
+
<div class="node">Lower memory</div><div class="arrow">+</div>
|
| 116 |
+
<div class="node">Potential speed gains</div><div class="arrow">+</div>
|
| 117 |
+
<div class="node">Trade-offs</div>
|
| 118 |
+
</div>
|
| 119 |
+
</div>
|
| 120 |
+
</section>
|
| 121 |
+
|
| 122 |
+
<section class="section">
|
| 123 |
+
<div class="eyebrow">02 · Precision</div>
|
| 124 |
+
<h2>Bits per parameter: the basic intuition</h2>
|
| 125 |
+
<p>The chart below shows <strong>theoretical raw weight storage</strong> relative to FP32. It ignores runtime overhead, metadata, KV cache, activations and mixed-precision components.</p>
|
| 126 |
+
<div class="barwrap">
|
| 127 |
+
<div class="bar" style="height:100%"><b>32</b><small>FP32</small></div>
|
| 128 |
+
<div class="bar" style="height:50%"><b>16</b><small>FP16/BF16</small></div>
|
| 129 |
+
<div class="bar" style="height:25%"><b>8</b><small>FP8</small></div>
|
| 130 |
+
<div class="bar" style="height:25%"><b>8</b><small>INT8</small></div>
|
| 131 |
+
<div class="bar" style="height:12.5%"><b>4</b><small>INT4</small></div>
|
| 132 |
+
</div>
|
| 133 |
+
<div class="callout warning"><strong>Lower bit width does not guarantee faster inference.</strong> Actual performance depends on kernels, hardware, memory bandwidth, runtime support and the quantization method.</div>
|
| 134 |
+
</section>
|
| 135 |
+
|
| 136 |
+
<section class="section">
|
| 137 |
+
<div class="eyebrow">03 · Memory estimator</div>
|
| 138 |
+
<h2>Estimate raw weight storage</h2>
|
| 139 |
+
<p>Use this simple calculator to estimate the theoretical storage of model weights at a chosen bit width.</p>
|
| 140 |
+
<div class="controls">
|
| 141 |
+
<div>
|
| 142 |
+
<label for="params">Model parameters (billions)</label>
|
| 143 |
+
<input id="params" type="number" min="0.1" step="0.1" value="8">
|
| 144 |
+
</div>
|
| 145 |
+
<div>
|
| 146 |
+
<label for="bits">Bits per parameter</label>
|
| 147 |
+
<select id="bits">
|
| 148 |
+
<option value="32">FP32 — 32 bit</option>
|
| 149 |
+
<option value="16" selected>FP16 / BF16 — 16 bit</option>
|
| 150 |
+
<option value="8">FP8 / INT8 — 8 bit</option>
|
| 151 |
+
<option value="4">INT4 — 4 bit</option>
|
| 152 |
+
<option value="2">2 bit — method dependent</option>
|
| 153 |
+
</select>
|
| 154 |
+
</div>
|
| 155 |
+
</div>
|
| 156 |
+
<div class="result">
|
| 157 |
+
<div class="big" id="memoryOut">16.00 GB</div>
|
| 158 |
+
<div class="muted">Approximate decimal GB for raw parameters only. Real deployment memory can be higher.</div>
|
| 159 |
+
</div>
|
| 160 |
+
</section>
|
| 161 |
+
|
| 162 |
+
<section class="section">
|
| 163 |
+
<div class="eyebrow">04 · Main approaches</div>
|
| 164 |
+
<h2>Quantization is not one technique</h2>
|
| 165 |
+
<div class="grid3">
|
| 166 |
+
<div class="card">
|
| 167 |
+
<h3>On-the-fly</h3>
|
| 168 |
+
<p class="muted">Quantize during model loading rather than distributing a separately pre-quantized checkpoint.</p>
|
| 169 |
+
<span class="pill">bitsandbytes</span>
|
| 170 |
+
</div>
|
| 171 |
+
<div class="card">
|
| 172 |
+
<h3>Post-training</h3>
|
| 173 |
+
<p class="muted">Quantize an already trained model, often using calibration or optimization to reduce error.</p>
|
| 174 |
+
<span class="pill">GPTQ</span><span class="pill">AWQ</span>
|
| 175 |
+
</div>
|
| 176 |
+
<div class="card">
|
| 177 |
+
<h3>Runtime ecosystem</h3>
|
| 178 |
+
<p class="muted">Convert and quantize for a deployment stack such as GGUF / llama.cpp.</p>
|
| 179 |
+
<span class="pill">GGUF</span><span class="pill">llama.cpp</span>
|
| 180 |
+
</div>
|
| 181 |
+
</div>
|
| 182 |
+
</section>
|
| 183 |
+
|
| 184 |
+
<section class="section">
|
| 185 |
+
<div class="eyebrow">05 · bitsandbytes</div>
|
| 186 |
+
<h2>4-bit and 8-bit loading in Transformers</h2>
|
| 187 |
+
<p>Hugging Face documents <strong>bitsandbytes</strong> as providing memory-efficient 8-bit and 4-bit linear layers and quantization integrations for Transformers.</p>
|
| 188 |
+
<div class="grid2">
|
| 189 |
+
<div class="card">
|
| 190 |
+
<h3>LLM.int8()</h3>
|
| 191 |
+
<p class="muted">An 8-bit method designed to preserve higher precision for sensitive computations instead of naively forcing everything into INT8.</p>
|
| 192 |
+
</div>
|
| 193 |
+
<div class="card">
|
| 194 |
+
<h3>QLoRA</h3>
|
| 195 |
+
<p class="muted">Uses 4-bit quantization with trainable low-rank adapter parameters, making parameter-efficient adaptation possible with a smaller memory footprint.</p>
|
| 196 |
+
</div>
|
| 197 |
+
</div>
|
| 198 |
+
<p><a href="https://huggingface.co/docs/transformers/en/quantization/bitsandbytes" target="_blank" rel="noopener">Hugging Face bitsandbytes documentation ↗</a></p>
|
| 199 |
+
</section>
|
| 200 |
+
|
| 201 |
+
<section class="section">
|
| 202 |
+
<div class="eyebrow">06 · GPTQ</div>
|
| 203 |
+
<h2>Error-aware post-training quantization</h2>
|
| 204 |
+
<p>Current Transformers documentation uses <strong>GPT-QModel</strong> as the maintained GPTQ backend. GPTQ is a post-training method that quantizes weight matrices while optimizing to reduce quantization error.</p>
|
| 205 |
+
<div class="callout">Hugging Face notes that current GPTQ workflows can quantize weights to low-bit representations such as INT4 and dequantize them during inference in optimized kernels.</div>
|
| 206 |
+
<p><a href="https://huggingface.co/docs/transformers/quantization/gptq" target="_blank" rel="noopener">Hugging Face GPTQ documentation ↗</a></p>
|
| 207 |
+
</section>
|
| 208 |
+
|
| 209 |
+
<section class="section">
|
| 210 |
+
<div class="eyebrow">07 · AWQ</div>
|
| 211 |
+
<h2>Activation-aware weight quantization</h2>
|
| 212 |
+
<p><strong>AWQ</strong> focuses on preserving weights that are especially important to model behavior while compressing the model to low-bit representations.</p>
|
| 213 |
+
<p>Transformers documents AWQ as an activation-aware approach designed for 4-bit compression with limited performance degradation.</p>
|
| 214 |
+
<p><a href="https://huggingface.co/docs/transformers/quantization/awq" target="_blank" rel="noopener">Hugging Face AWQ documentation ↗</a></p>
|
| 215 |
+
</section>
|
| 216 |
+
|
| 217 |
+
<section class="section">
|
| 218 |
+
<div class="eyebrow">08 · GGUF / llama.cpp</div>
|
| 219 |
+
<h2>Quantization for local and portable inference</h2>
|
| 220 |
+
<p>The llama.cpp ecosystem provides many quantized GGUF tensor types and tooling for converting higher-precision GGUF models into smaller quantized variants.</p>
|
| 221 |
+
<div class="code">High-precision model
|
| 222 |
+
↓
|
| 223 |
+
Convert to GGUF
|
| 224 |
+
↓
|
| 225 |
+
llama-quantize
|
| 226 |
+
↓
|
| 227 |
+
Q8 / Q6 / Q5 / Q4 / lower-bit variants
|
| 228 |
+
↓
|
| 229 |
+
Evaluate quality + performance
|
| 230 |
+
↓
|
| 231 |
+
Run with llama.cpp</div>
|
| 232 |
+
<p class="muted">llama.cpp documents integer quantization from very low bit widths through 8-bit variants. The practical choice depends on model family, quality target and hardware.</p>
|
| 233 |
+
<p><a href="https://github.com/ggml-org/llama.cpp/tree/master/tools/quantize" target="_blank" rel="noopener">llama.cpp quantization tools ↗</a></p>
|
| 234 |
+
</section>
|
| 235 |
+
|
| 236 |
+
<section class="section">
|
| 237 |
+
<div class="eyebrow">09 · Comparison</div>
|
| 238 |
+
<h2>Common quantization directions</h2>
|
| 239 |
+
<div class="tablewrap">
|
| 240 |
+
<table>
|
| 241 |
+
<thead><tr><th>Approach</th><th>Typical bit width</th><th>Strength</th><th>Watch for</th></tr></thead>
|
| 242 |
+
<tbody>
|
| 243 |
+
<tr><td>bitsandbytes</td><td>4 / 8</td><td>Convenient Transformers integration and on-the-fly loading</td><td>Hardware/backend support and training limitations</td></tr>
|
| 244 |
+
<tr><td>GPTQ</td><td>Commonly 4; other bit widths supported by current backends</td><td>Post-training compression with error-aware optimization</td><td>Kernel, model and checkpoint compatibility</td></tr>
|
| 245 |
+
<tr><td>AWQ</td><td>4</td><td>Activation-aware preservation of important weights</td><td>Toolchain and runtime compatibility</td></tr>
|
| 246 |
+
<tr><td>GGUF / llama.cpp</td><td>Multiple low-bit types</td><td>Strong local-inference ecosystem and many quantization variants</td><td>Model architecture support and quality/runtime trade-offs</td></tr>
|
| 247 |
+
<tr><td>FP8</td><td>8</td><td>Lower precision while remaining floating point</td><td>Hardware and kernel support</td></tr>
|
| 248 |
+
</tbody>
|
| 249 |
+
</table>
|
| 250 |
+
</div>
|
| 251 |
+
</section>
|
| 252 |
+
|
| 253 |
+
<section class="section">
|
| 254 |
+
<div class="eyebrow">10 · Decision helper</div>
|
| 255 |
+
<h2>Which direction should you investigate?</h2>
|
| 256 |
+
<p>Select a deployment goal. This is a starting point, not a universal recommendation.</p>
|
| 257 |
+
<div class="tabs">
|
| 258 |
+
<button class="tab active" data-answer="hf">Transformers simplicity</button>
|
| 259 |
+
<button class="tab" data-answer="local">Local / consumer hardware</button>
|
| 260 |
+
<button class="tab" data-answer="gpu">GPU server inference</button>
|
| 261 |
+
<button class="tab" data-answer="tune">Fine-tuning</button>
|
| 262 |
+
</div>
|
| 263 |
+
<div class="answer" id="answer">
|
| 264 |
+
<strong>Start by evaluating bitsandbytes.</strong>
|
| 265 |
+
<p class="muted">Its 4-bit and 8-bit Transformers integration makes it a practical entry point when your model and hardware are supported.</p>
|
| 266 |
+
</div>
|
| 267 |
+
</section>
|
| 268 |
+
|
| 269 |
+
<section class="section">
|
| 270 |
+
<div class="eyebrow">11 · Trade-offs</div>
|
| 271 |
+
<h2>What should you measure?</h2>
|
| 272 |
+
<div class="grid4">
|
| 273 |
+
<div class="card"><h3>Memory</h3><p class="muted">How much RAM or VRAM is actually required?</p></div>
|
| 274 |
+
<div class="card"><h3>Quality</h3><p class="muted">How much task performance changes after quantization?</p></div>
|
| 275 |
+
<div class="card"><h3>Latency</h3><p class="muted">Does the runtime and hardware actually become faster?</p></div>
|
| 276 |
+
<div class="card"><h3>Compatibility</h3><p class="muted">Can your serving stack load and accelerate the chosen format?</p></div>
|
| 277 |
+
</div>
|
| 278 |
+
<div class="callout warning"><strong>Always benchmark on the real workload.</strong> A smaller checkpoint can still perform worse operationally if the runtime lacks optimized kernels for that quantization.</div>
|
| 279 |
+
</section>
|
| 280 |
+
|
| 281 |
+
<section class="section">
|
| 282 |
+
<div class="eyebrow">12 · Common mistakes</div>
|
| 283 |
+
<h2>Quantization misconceptions</h2>
|
| 284 |
+
<div class="grid3">
|
| 285 |
+
<div class="card"><h3>Bits ≠ method</h3><p class="muted">Two 4-bit methods can behave very differently.</p></div>
|
| 286 |
+
<div class="card"><h3>Smaller ≠ faster</h3><p class="muted">Speed depends on kernels, hardware and runtime support.</p></div>
|
| 287 |
+
<div class="card"><h3>Format ≠ quantization</h3><p class="muted">GGUF or Safetensors are serialization formats; quantization describes numerical representation and method.</p></div>
|
| 288 |
+
<div class="card"><h3>Memory ≠ file size only</h3><p class="muted">KV cache, activations and runtime overhead also matter.</p></div>
|
| 289 |
+
<div class="card"><h3>Quality loss is task-specific</h3><p class="muted">Benchmark the model on the tasks that matter to you.</p></div>
|
| 290 |
+
<div class="card"><h3>Support changes</h3><p class="muted">Quantization libraries and hardware backends evolve quickly.</p></div>
|
| 291 |
+
</div>
|
| 292 |
+
</section>
|
| 293 |
+
|
| 294 |
+
<section class="section">
|
| 295 |
+
<div class="eyebrow">13 · Quick checklist</div>
|
| 296 |
+
<h2>Before choosing a quantization</h2>
|
| 297 |
+
<div class="tablewrap">
|
| 298 |
+
<table>
|
| 299 |
+
<thead><tr><th>Question</th><th>Why it matters</th></tr></thead>
|
| 300 |
+
<tbody>
|
| 301 |
+
<tr><td>What hardware will run the model?</td><td>Backend support and optimized kernels differ by platform.</td></tr>
|
| 302 |
+
<tr><td>Which runtime will serve it?</td><td>Not every runtime supports every quantization method.</td></tr>
|
| 303 |
+
<tr><td>What memory limit do you have?</td><td>Defines how aggressive compression may need to be.</td></tr>
|
| 304 |
+
<tr><td>What quality loss is acceptable?</td><td>Lower bit widths can affect downstream performance.</td></tr>
|
| 305 |
+
<tr><td>Do you need fine-tuning?</td><td>Some workflows support PEFT or adapter training better than others.</td></tr>
|
| 306 |
+
<tr><td>Do you need portability?</td><td>A highly optimized method may tie you to a specific runtime or hardware stack.</td></tr>
|
| 307 |
+
</tbody>
|
| 308 |
+
</table>
|
| 309 |
+
</div>
|
| 310 |
+
</section>
|
| 311 |
+
|
| 312 |
+
<section class="section">
|
| 313 |
+
<div class="eyebrow">Next</div>
|
| 314 |
+
<h2>Continue the Open Weight series</h2>
|
| 315 |
+
<div class="grid3">
|
| 316 |
+
<div class="card"><h3>Open Weight Explorer</h3><p class="muted">Understand tensors, model weights and the deployment stack.</p></div>
|
| 317 |
+
<div class="card"><h3>Weight Format Explorer</h3><p class="muted">Safetensors, GGUF, metadata, sharding and conversion.</p></div>
|
| 318 |
+
<div class="card"><h3>Model Portability Explorer</h3><p class="muted">Formats, runtimes, hardware and compatibility.</p></div>
|
| 319 |
+
</div>
|
| 320 |
+
</section>
|
| 321 |
+
|
| 322 |
+
<section class="section">
|
| 323 |
+
<div class="eyebrow">Primary sources</div>
|
| 324 |
+
<h2>Technical references</h2>
|
| 325 |
+
<p><a href="https://huggingface.co/docs/transformers/quantization/overview" target="_blank" rel="noopener">Hugging Face Transformers — Quantization overview ↗</a></p>
|
| 326 |
+
<p><a href="https://huggingface.co/docs/transformers/en/quantization/bitsandbytes" target="_blank" rel="noopener">Hugging Face — bitsandbytes ↗</a></p>
|
| 327 |
+
<p><a href="https://huggingface.co/docs/transformers/quantization/gptq" target="_blank" rel="noopener">Hugging Face — GPTQ ↗</a></p>
|
| 328 |
+
<p><a href="https://huggingface.co/docs/transformers/quantization/awq" target="_blank" rel="noopener">Hugging Face — AWQ ↗</a></p>
|
| 329 |
+
<p><a href="https://github.com/ggml-org/llama.cpp/tree/master/tools/quantize" target="_blank" rel="noopener">llama.cpp — quantization tools ↗</a></p>
|
| 330 |
+
</section>
|
| 331 |
+
|
| 332 |
+
<footer>
|
| 333 |
+
<strong style="color:white">Open Weight</strong><br>
|
| 334 |
+
Open weights. Portable models. Deployable AI.<br><br>
|
| 335 |
+
Collaboration: open-weight AI, model infrastructure, inference, deployment, research and ecosystem partnerships.<br>
|
| 336 |
+
Contact: <a href="mailto:agenten@magenta.de">agenten@magenta.de</a>
|
| 337 |
+
</footer>
|
| 338 |
+
</div>
|
| 339 |
+
|
| 340 |
+
<script>
|
| 341 |
+
function updateMemory(){
|
| 342 |
+
const p = Math.max(0, parseFloat(document.getElementById('params').value)||0);
|
| 343 |
+
const b = Math.max(0, parseFloat(document.getElementById('bits').value)||0);
|
| 344 |
+
const gb = p * b / 8;
|
| 345 |
+
document.getElementById('memoryOut').textContent = gb.toFixed(2) + ' GB';
|
| 346 |
+
}
|
| 347 |
+
document.getElementById('params').addEventListener('input',updateMemory);
|
| 348 |
+
document.getElementById('bits').addEventListener('change',updateMemory);
|
| 349 |
+
|
| 350 |
+
const answers={
|
| 351 |
+
hf:`<strong>Start by evaluating bitsandbytes.</strong><p class="muted">Its 4-bit and 8-bit Transformers integration makes it a practical entry point when your model and hardware are supported.</p>`,
|
| 352 |
+
local:`<strong>Investigate GGUF / llama.cpp quantization.</strong><p class="muted">The ecosystem offers many low-bit formats for local inference across supported CPUs, GPUs and Apple Silicon workflows.</p>`,
|
| 353 |
+
gpu:`<strong>Choose the serving runtime first.</strong><p class="muted">Then compare the methods it accelerates well — such as supported FP8, GPTQ, AWQ, bitsandbytes or other native quantization paths.</p>`,
|
| 354 |
+
tune:`<strong>Look closely at 4-bit PEFT / QLoRA workflows.</strong><p class="muted">bitsandbytes is a common entry point because it combines low-bit loading with trainable adapter parameters.</p>`
|
| 355 |
+
};
|
| 356 |
+
document.querySelectorAll('.tab').forEach(btn=>{
|
| 357 |
+
btn.addEventListener('click',()=>{
|
| 358 |
+
document.querySelectorAll('.tab').forEach(x=>x.classList.remove('active'));
|
| 359 |
+
btn.classList.add('active');
|
| 360 |
+
document.getElementById('answer').innerHTML=answers[btn.dataset.answer];
|
| 361 |
+
});
|
| 362 |
+
});
|
| 363 |
+
updateMemory();
|
| 364 |
+
</script>
|
| 365 |
+
</body>
|
| 366 |
+
</html>
|