Agenten commited on
Commit
762c004
·
verified ·
1 Parent(s): e41dc95

Upload 2 files

Browse files
Files changed (2) hide show
  1. README.md +71 -0
  2. index.html +366 -0
README.md ADDED
@@ -0,0 +1,71 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: Quantization Explorer
3
+ emoji: ⚙️
4
+ colorFrom: blue
5
+ colorTo: indigo
6
+ sdk: static
7
+ pinned: false
8
+ license: mit
9
+ short_description: "Explore quantization: FP8, INT8, INT4 and trade-offs."
10
+ ---
11
+
12
+ # Quantization Explorer
13
+
14
+ **Quantization Explorer** is an educational Hugging Face Space by the [`open-weight`](https://huggingface.co/open-weight) organization.
15
+
16
+ It explains how model quantization reduces memory requirements by representing weights at lower precision, and how common approaches such as **FP8, INT8, INT4, bitsandbytes, GPTQ, AWQ and GGUF quantization** differ in purpose and trade-offs.
17
+
18
+ ## What you can explore
19
+
20
+ - What model quantization is
21
+ - FP16/BF16 vs. FP8 vs. INT8 vs. INT4
22
+ - Theoretical raw weight memory
23
+ - Post-training quantization
24
+ - On-the-fly quantization
25
+ - Calibration-based methods
26
+ - bitsandbytes
27
+ - GPTQ
28
+ - AWQ
29
+ - GGUF / llama.cpp quantization
30
+ - Quality, speed and compatibility trade-offs
31
+ - A simple quantization decision helper
32
+
33
+ ## Core idea
34
+
35
+ ```text
36
+ Higher-precision weights
37
+ ↓
38
+ Quantization method
39
+ ↓
40
+ Lower-bit representation
41
+ ↓
42
+ Lower memory / storage
43
+ ↓
44
+ Potential speed benefits
45
+ +
46
+ Possible quality / compatibility trade-offs
47
+ ```
48
+
49
+ ## Primary references
50
+
51
+ - Hugging Face Transformers — Quantization overview: https://huggingface.co/docs/transformers/quantization/overview
52
+ - bitsandbytes: https://huggingface.co/docs/transformers/en/quantization/bitsandbytes
53
+ - GPTQ: https://huggingface.co/docs/transformers/quantization/gptq
54
+ - AWQ: https://huggingface.co/docs/transformers/quantization/awq
55
+ - llama.cpp quantization: https://github.com/ggml-org/llama.cpp/tree/master/tools/quantize
56
+
57
+ ## Related organization
58
+
59
+ Open Weight
60
+ https://huggingface.co/open-weight
61
+
62
+ ## Related project
63
+
64
+ Open Weights
65
+ https://huggingface.co/open-weights
66
+
67
+ ## Collaboration
68
+
69
+ Open-weight AI, model infrastructure, inference, deployment, research and ecosystem partnerships.
70
+
71
+ **Contact:** agenten@magenta.de
index.html ADDED
@@ -0,0 +1,366 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ <!doctype html>
2
+ <html lang="en">
3
+ <head>
4
+ <meta charset="utf-8">
5
+ <meta name="viewport" content="width=device-width,initial-scale=1">
6
+ <meta name="description" content="Explore AI model quantization: FP16, BF16, FP8, INT8, INT4, bitsandbytes, GPTQ, AWQ and GGUF. Compare memory, quality and deployment trade-offs.">
7
+ <meta name="theme-color" content="#07111f">
8
+ <title>Quantization Explorer — FP8, INT8, INT4 & Open-Weight Deployment</title>
9
+ <style>
10
+ :root{
11
+ --bg:#07111f;--panel:#0d1b2d;--panel2:#10233a;--text:#edf7ff;--muted:#9fb4c8;
12
+ --line:#24445f;--cyan:#55d9ff;--blue:#6b8cff;--green:#79f2c0;--gold:#ffd580;
13
+ --red:#ff9e9e;--shadow:0 18px 60px rgba(0,0,0,.25)
14
+ }
15
+ *{box-sizing:border-box}
16
+ html{scroll-behavior:smooth}
17
+ body{
18
+ margin:0;color:var(--text);
19
+ font:16px/1.65 Inter,ui-sans-serif,system-ui,-apple-system,BlinkMacSystemFont,"Segoe UI",sans-serif;
20
+ background:
21
+ radial-gradient(circle at 12% 0%,rgba(85,217,255,.12),transparent 29%),
22
+ radial-gradient(circle at 90% 12%,rgba(107,140,255,.13),transparent 26%),
23
+ var(--bg)
24
+ }
25
+ a{color:var(--cyan);text-decoration:none}
26
+ a:hover{text-decoration:underline}
27
+ .wrap{max-width:1180px;margin:auto;padding:0 22px}
28
+ .hero{padding:72px 0 35px}
29
+ .badge{display:inline-flex;padding:7px 12px;border:1px solid var(--line);border-radius:999px;background:rgba(13,27,45,.75);color:#c8efff;font-size:14px}
30
+ h1{font-size:clamp(42px,7vw,78px);line-height:1;letter-spacing:-.055em;margin:20px 0;max-width:980px}
31
+ .gradient{background:linear-gradient(90deg,var(--cyan),#b6c3ff);-webkit-background-clip:text;background-clip:text;color:transparent}
32
+ .lead{font-size:clamp(18px,2.2vw,24px);max-width:900px;color:#cbdbe8;margin:0 0 28px}
33
+ .cta{display:flex;gap:12px;flex-wrap:wrap}
34
+ .btn{display:inline-block;padding:11px 16px;border:1px solid var(--line);border-radius:12px;font-weight:750}
35
+ .btn.primary{border:0;color:white;background:linear-gradient(135deg,#157aa8,#5269df)}
36
+ .quick{display:grid;grid-template-columns:repeat(4,1fr);gap:14px;margin:32px 0 56px}
37
+ .card,.section{border:1px solid var(--line);background:linear-gradient(180deg,rgba(16,35,58,.93),rgba(10,25,42,.93));border-radius:20px;box-shadow:var(--shadow)}
38
+ .card{padding:18px}
39
+ .card strong{display:block;font-size:23px;color:#fff}
40
+ .card span,.muted{color:var(--muted)}
41
+ .section{padding:28px;margin:22px 0}
42
+ .eyebrow{text-transform:uppercase;letter-spacing:.14em;font-size:12px;color:var(--cyan);font-weight:850}
43
+ h2{font-size:clamp(28px,4vw,45px);letter-spacing:-.03em;margin:6px 0 10px}
44
+ h3{font-size:22px;margin:0 0 8px}
45
+ .grid2{display:grid;grid-template-columns:1fr 1fr;gap:18px}
46
+ .grid3{display:grid;grid-template-columns:repeat(3,1fr);gap:14px}
47
+ .grid4{display:grid;grid-template-columns:repeat(4,1fr);gap:14px}
48
+ .callout{padding:15px 17px;border-left:4px solid var(--cyan);border-radius:9px;background:rgba(85,217,255,.06);margin:18px 0}
49
+ .warning{border-left-color:var(--gold);background:rgba(255,213,128,.06)}
50
+ .diagram{padding:22px;border:1px solid #2d5876;border-radius:17px;background:#091829;overflow:auto;margin:20px 0}
51
+ .flow{display:flex;gap:9px;align-items:center;min-width:900px}
52
+ .node{min-width:135px;padding:14px 12px;text-align:center;border:1px solid #34617d;background:#102842;border-radius:13px;font-weight:800}
53
+ .arrow{font-size:24px;color:var(--cyan)}
54
+ .tablewrap{overflow:auto}
55
+ table{width:100%;border-collapse:collapse;min-width:760px}
56
+ th,td{padding:13px;border-bottom:1px solid #23435e;text-align:left;vertical-align:top}
57
+ th{font-size:12px;text-transform:uppercase;letter-spacing:.07em;color:#c9efff}
58
+ .pill{display:inline-block;padding:5px 9px;margin:3px;border:1px solid #315b78;border-radius:999px;background:#102842;color:#c8ecff;font-size:13px}
59
+ .code{white-space:pre-wrap;padding:17px;border:1px solid #223f58;border-radius:14px;background:#06101c;color:#bfeeff;font:14px/1.6 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace;overflow:auto}
60
+ .barwrap{display:grid;grid-template-columns:repeat(5,1fr);gap:12px;align-items:end;height:220px;margin:28px 0 42px}
61
+ .bar{position:relative;border-radius:12px 12px 4px 4px;background:linear-gradient(180deg,var(--cyan),#5368de);min-height:28px}
62
+ .bar b{position:absolute;top:10px;left:0;right:0;text-align:center;color:#06111f}
63
+ .bar small{position:absolute;bottom:-30px;left:0;right:0;text-align:center;color:var(--muted)}
64
+ .controls{display:grid;grid-template-columns:1fr 1fr;gap:16px}
65
+ label{display:block;color:#c8efff;font-weight:700;margin-bottom:7px}
66
+ input,select,button{
67
+ width:100%;background:#0a1c2f;color:var(--text);border:1px solid #315b78;border-radius:11px;padding:11px 12px;font:inherit
68
+ }
69
+ .result{margin-top:18px;padding:18px;border:1px solid #315b78;background:#091a2b;border-radius:15px}
70
+ .big{font-size:34px;font-weight:850;color:white}
71
+ .tabs{display:flex;gap:9px;flex-wrap:wrap;margin:16px 0}
72
+ .tab{width:auto;cursor:pointer;border:1px solid #315b78;background:#0c2035;color:#d9f2ff;border-radius:999px;padding:8px 12px;font-weight:700}
73
+ .tab.active{background:linear-gradient(135deg,#177aa6,#4e66d8);border-color:transparent}
74
+ .answer{padding:18px;border:1px solid #315b78;border-radius:15px;background:#0a1c2f;min-height:112px}
75
+ .good{color:var(--green);font-weight:800}
76
+ .caution{color:var(--gold);font-weight:800}
77
+ .bad{color:var(--red);font-weight:800}
78
+ footer{padding:50px 0 68px;color:var(--muted)}
79
+ @media(max-width:820px){
80
+ .quick,.grid2,.grid3,.grid4,.controls{grid-template-columns:1fr}
81
+ .hero{padding-top:48px}.section{padding:20px}
82
+ .barwrap{grid-template-columns:repeat(5,minmax(52px,1fr))}
83
+ }
84
+ </style>
85
+ </head>
86
+ <body>
87
+ <div class="wrap">
88
+ <header class="hero">
89
+ <div class="badge">Open Weight · Quantization Explorer</div>
90
+ <h1>Make models smaller. Understand the <span class="gradient">trade-offs.</span></h1>
91
+ <p class="lead">A practical guide to model quantization — from FP16 and BF16 to FP8, INT8, INT4, bitsandbytes, GPTQ, AWQ and GGUF-based local inference.</p>
92
+ <div class="cta">
93
+ <a class="btn primary" href="#basics">Start exploring</a>
94
+ <a class="btn" href="https://huggingface.co/open-weight" target="_blank" rel="noopener">Open Weight organization ↗</a>
95
+ </div>
96
+ </header>
97
+
98
+ <div class="quick">
99
+ <div class="card"><strong>FP8</strong><span>8-bit floating-point workflows</span></div>
100
+ <div class="card"><strong>INT8</strong><span>Lower-memory integer inference</span></div>
101
+ <div class="card"><strong>INT4</strong><span>High compression for deployment</span></div>
102
+ <div class="card"><strong>Trade-offs</strong><span>Memory · quality · speed · support</span></div>
103
+ </div>
104
+
105
+ <section class="section" id="basics">
106
+ <div class="eyebrow">01 · Foundation</div>
107
+ <h2>What is model quantization?</h2>
108
+ <p><strong>Quantization reduces the precision used to represent model weights or activations.</strong> The goal is usually to reduce memory requirements and make models easier or cheaper to run while preserving as much model quality as possible.</p>
109
+ <div class="callout">Hugging Face describes quantization as lowering model memory requirements by storing weights at lower precision while trying to preserve accuracy.</div>
110
+ <div class="diagram">
111
+ <div class="flow">
112
+ <div class="node">FP32 / BF16 / FP16</div><div class="arrow">→</div>
113
+ <div class="node">Quantization method</div><div class="arrow">→</div>
114
+ <div class="node">FP8 / INT8 / INT4</div><div class="arrow">→</div>
115
+ <div class="node">Lower memory</div><div class="arrow">+</div>
116
+ <div class="node">Potential speed gains</div><div class="arrow">+</div>
117
+ <div class="node">Trade-offs</div>
118
+ </div>
119
+ </div>
120
+ </section>
121
+
122
+ <section class="section">
123
+ <div class="eyebrow">02 · Precision</div>
124
+ <h2>Bits per parameter: the basic intuition</h2>
125
+ <p>The chart below shows <strong>theoretical raw weight storage</strong> relative to FP32. It ignores runtime overhead, metadata, KV cache, activations and mixed-precision components.</p>
126
+ <div class="barwrap">
127
+ <div class="bar" style="height:100%"><b>32</b><small>FP32</small></div>
128
+ <div class="bar" style="height:50%"><b>16</b><small>FP16/BF16</small></div>
129
+ <div class="bar" style="height:25%"><b>8</b><small>FP8</small></div>
130
+ <div class="bar" style="height:25%"><b>8</b><small>INT8</small></div>
131
+ <div class="bar" style="height:12.5%"><b>4</b><small>INT4</small></div>
132
+ </div>
133
+ <div class="callout warning"><strong>Lower bit width does not guarantee faster inference.</strong> Actual performance depends on kernels, hardware, memory bandwidth, runtime support and the quantization method.</div>
134
+ </section>
135
+
136
+ <section class="section">
137
+ <div class="eyebrow">03 · Memory estimator</div>
138
+ <h2>Estimate raw weight storage</h2>
139
+ <p>Use this simple calculator to estimate the theoretical storage of model weights at a chosen bit width.</p>
140
+ <div class="controls">
141
+ <div>
142
+ <label for="params">Model parameters (billions)</label>
143
+ <input id="params" type="number" min="0.1" step="0.1" value="8">
144
+ </div>
145
+ <div>
146
+ <label for="bits">Bits per parameter</label>
147
+ <select id="bits">
148
+ <option value="32">FP32 — 32 bit</option>
149
+ <option value="16" selected>FP16 / BF16 — 16 bit</option>
150
+ <option value="8">FP8 / INT8 — 8 bit</option>
151
+ <option value="4">INT4 — 4 bit</option>
152
+ <option value="2">2 bit — method dependent</option>
153
+ </select>
154
+ </div>
155
+ </div>
156
+ <div class="result">
157
+ <div class="big" id="memoryOut">16.00 GB</div>
158
+ <div class="muted">Approximate decimal GB for raw parameters only. Real deployment memory can be higher.</div>
159
+ </div>
160
+ </section>
161
+
162
+ <section class="section">
163
+ <div class="eyebrow">04 · Main approaches</div>
164
+ <h2>Quantization is not one technique</h2>
165
+ <div class="grid3">
166
+ <div class="card">
167
+ <h3>On-the-fly</h3>
168
+ <p class="muted">Quantize during model loading rather than distributing a separately pre-quantized checkpoint.</p>
169
+ <span class="pill">bitsandbytes</span>
170
+ </div>
171
+ <div class="card">
172
+ <h3>Post-training</h3>
173
+ <p class="muted">Quantize an already trained model, often using calibration or optimization to reduce error.</p>
174
+ <span class="pill">GPTQ</span><span class="pill">AWQ</span>
175
+ </div>
176
+ <div class="card">
177
+ <h3>Runtime ecosystem</h3>
178
+ <p class="muted">Convert and quantize for a deployment stack such as GGUF / llama.cpp.</p>
179
+ <span class="pill">GGUF</span><span class="pill">llama.cpp</span>
180
+ </div>
181
+ </div>
182
+ </section>
183
+
184
+ <section class="section">
185
+ <div class="eyebrow">05 · bitsandbytes</div>
186
+ <h2>4-bit and 8-bit loading in Transformers</h2>
187
+ <p>Hugging Face documents <strong>bitsandbytes</strong> as providing memory-efficient 8-bit and 4-bit linear layers and quantization integrations for Transformers.</p>
188
+ <div class="grid2">
189
+ <div class="card">
190
+ <h3>LLM.int8()</h3>
191
+ <p class="muted">An 8-bit method designed to preserve higher precision for sensitive computations instead of naively forcing everything into INT8.</p>
192
+ </div>
193
+ <div class="card">
194
+ <h3>QLoRA</h3>
195
+ <p class="muted">Uses 4-bit quantization with trainable low-rank adapter parameters, making parameter-efficient adaptation possible with a smaller memory footprint.</p>
196
+ </div>
197
+ </div>
198
+ <p><a href="https://huggingface.co/docs/transformers/en/quantization/bitsandbytes" target="_blank" rel="noopener">Hugging Face bitsandbytes documentation ↗</a></p>
199
+ </section>
200
+
201
+ <section class="section">
202
+ <div class="eyebrow">06 · GPTQ</div>
203
+ <h2>Error-aware post-training quantization</h2>
204
+ <p>Current Transformers documentation uses <strong>GPT-QModel</strong> as the maintained GPTQ backend. GPTQ is a post-training method that quantizes weight matrices while optimizing to reduce quantization error.</p>
205
+ <div class="callout">Hugging Face notes that current GPTQ workflows can quantize weights to low-bit representations such as INT4 and dequantize them during inference in optimized kernels.</div>
206
+ <p><a href="https://huggingface.co/docs/transformers/quantization/gptq" target="_blank" rel="noopener">Hugging Face GPTQ documentation ↗</a></p>
207
+ </section>
208
+
209
+ <section class="section">
210
+ <div class="eyebrow">07 · AWQ</div>
211
+ <h2>Activation-aware weight quantization</h2>
212
+ <p><strong>AWQ</strong> focuses on preserving weights that are especially important to model behavior while compressing the model to low-bit representations.</p>
213
+ <p>Transformers documents AWQ as an activation-aware approach designed for 4-bit compression with limited performance degradation.</p>
214
+ <p><a href="https://huggingface.co/docs/transformers/quantization/awq" target="_blank" rel="noopener">Hugging Face AWQ documentation ↗</a></p>
215
+ </section>
216
+
217
+ <section class="section">
218
+ <div class="eyebrow">08 · GGUF / llama.cpp</div>
219
+ <h2>Quantization for local and portable inference</h2>
220
+ <p>The llama.cpp ecosystem provides many quantized GGUF tensor types and tooling for converting higher-precision GGUF models into smaller quantized variants.</p>
221
+ <div class="code">High-precision model
222
+ ↓
223
+ Convert to GGUF
224
+ ↓
225
+ llama-quantize
226
+ ↓
227
+ Q8 / Q6 / Q5 / Q4 / lower-bit variants
228
+ ↓
229
+ Evaluate quality + performance
230
+ ↓
231
+ Run with llama.cpp</div>
232
+ <p class="muted">llama.cpp documents integer quantization from very low bit widths through 8-bit variants. The practical choice depends on model family, quality target and hardware.</p>
233
+ <p><a href="https://github.com/ggml-org/llama.cpp/tree/master/tools/quantize" target="_blank" rel="noopener">llama.cpp quantization tools ↗</a></p>
234
+ </section>
235
+
236
+ <section class="section">
237
+ <div class="eyebrow">09 · Comparison</div>
238
+ <h2>Common quantization directions</h2>
239
+ <div class="tablewrap">
240
+ <table>
241
+ <thead><tr><th>Approach</th><th>Typical bit width</th><th>Strength</th><th>Watch for</th></tr></thead>
242
+ <tbody>
243
+ <tr><td>bitsandbytes</td><td>4 / 8</td><td>Convenient Transformers integration and on-the-fly loading</td><td>Hardware/backend support and training limitations</td></tr>
244
+ <tr><td>GPTQ</td><td>Commonly 4; other bit widths supported by current backends</td><td>Post-training compression with error-aware optimization</td><td>Kernel, model and checkpoint compatibility</td></tr>
245
+ <tr><td>AWQ</td><td>4</td><td>Activation-aware preservation of important weights</td><td>Toolchain and runtime compatibility</td></tr>
246
+ <tr><td>GGUF / llama.cpp</td><td>Multiple low-bit types</td><td>Strong local-inference ecosystem and many quantization variants</td><td>Model architecture support and quality/runtime trade-offs</td></tr>
247
+ <tr><td>FP8</td><td>8</td><td>Lower precision while remaining floating point</td><td>Hardware and kernel support</td></tr>
248
+ </tbody>
249
+ </table>
250
+ </div>
251
+ </section>
252
+
253
+ <section class="section">
254
+ <div class="eyebrow">10 · Decision helper</div>
255
+ <h2>Which direction should you investigate?</h2>
256
+ <p>Select a deployment goal. This is a starting point, not a universal recommendation.</p>
257
+ <div class="tabs">
258
+ <button class="tab active" data-answer="hf">Transformers simplicity</button>
259
+ <button class="tab" data-answer="local">Local / consumer hardware</button>
260
+ <button class="tab" data-answer="gpu">GPU server inference</button>
261
+ <button class="tab" data-answer="tune">Fine-tuning</button>
262
+ </div>
263
+ <div class="answer" id="answer">
264
+ <strong>Start by evaluating bitsandbytes.</strong>
265
+ <p class="muted">Its 4-bit and 8-bit Transformers integration makes it a practical entry point when your model and hardware are supported.</p>
266
+ </div>
267
+ </section>
268
+
269
+ <section class="section">
270
+ <div class="eyebrow">11 · Trade-offs</div>
271
+ <h2>What should you measure?</h2>
272
+ <div class="grid4">
273
+ <div class="card"><h3>Memory</h3><p class="muted">How much RAM or VRAM is actually required?</p></div>
274
+ <div class="card"><h3>Quality</h3><p class="muted">How much task performance changes after quantization?</p></div>
275
+ <div class="card"><h3>Latency</h3><p class="muted">Does the runtime and hardware actually become faster?</p></div>
276
+ <div class="card"><h3>Compatibility</h3><p class="muted">Can your serving stack load and accelerate the chosen format?</p></div>
277
+ </div>
278
+ <div class="callout warning"><strong>Always benchmark on the real workload.</strong> A smaller checkpoint can still perform worse operationally if the runtime lacks optimized kernels for that quantization.</div>
279
+ </section>
280
+
281
+ <section class="section">
282
+ <div class="eyebrow">12 · Common mistakes</div>
283
+ <h2>Quantization misconceptions</h2>
284
+ <div class="grid3">
285
+ <div class="card"><h3>Bits ≠ method</h3><p class="muted">Two 4-bit methods can behave very differently.</p></div>
286
+ <div class="card"><h3>Smaller ≠ faster</h3><p class="muted">Speed depends on kernels, hardware and runtime support.</p></div>
287
+ <div class="card"><h3>Format ≠ quantization</h3><p class="muted">GGUF or Safetensors are serialization formats; quantization describes numerical representation and method.</p></div>
288
+ <div class="card"><h3>Memory ≠ file size only</h3><p class="muted">KV cache, activations and runtime overhead also matter.</p></div>
289
+ <div class="card"><h3>Quality loss is task-specific</h3><p class="muted">Benchmark the model on the tasks that matter to you.</p></div>
290
+ <div class="card"><h3>Support changes</h3><p class="muted">Quantization libraries and hardware backends evolve quickly.</p></div>
291
+ </div>
292
+ </section>
293
+
294
+ <section class="section">
295
+ <div class="eyebrow">13 · Quick checklist</div>
296
+ <h2>Before choosing a quantization</h2>
297
+ <div class="tablewrap">
298
+ <table>
299
+ <thead><tr><th>Question</th><th>Why it matters</th></tr></thead>
300
+ <tbody>
301
+ <tr><td>What hardware will run the model?</td><td>Backend support and optimized kernels differ by platform.</td></tr>
302
+ <tr><td>Which runtime will serve it?</td><td>Not every runtime supports every quantization method.</td></tr>
303
+ <tr><td>What memory limit do you have?</td><td>Defines how aggressive compression may need to be.</td></tr>
304
+ <tr><td>What quality loss is acceptable?</td><td>Lower bit widths can affect downstream performance.</td></tr>
305
+ <tr><td>Do you need fine-tuning?</td><td>Some workflows support PEFT or adapter training better than others.</td></tr>
306
+ <tr><td>Do you need portability?</td><td>A highly optimized method may tie you to a specific runtime or hardware stack.</td></tr>
307
+ </tbody>
308
+ </table>
309
+ </div>
310
+ </section>
311
+
312
+ <section class="section">
313
+ <div class="eyebrow">Next</div>
314
+ <h2>Continue the Open Weight series</h2>
315
+ <div class="grid3">
316
+ <div class="card"><h3>Open Weight Explorer</h3><p class="muted">Understand tensors, model weights and the deployment stack.</p></div>
317
+ <div class="card"><h3>Weight Format Explorer</h3><p class="muted">Safetensors, GGUF, metadata, sharding and conversion.</p></div>
318
+ <div class="card"><h3>Model Portability Explorer</h3><p class="muted">Formats, runtimes, hardware and compatibility.</p></div>
319
+ </div>
320
+ </section>
321
+
322
+ <section class="section">
323
+ <div class="eyebrow">Primary sources</div>
324
+ <h2>Technical references</h2>
325
+ <p><a href="https://huggingface.co/docs/transformers/quantization/overview" target="_blank" rel="noopener">Hugging Face Transformers — Quantization overview ↗</a></p>
326
+ <p><a href="https://huggingface.co/docs/transformers/en/quantization/bitsandbytes" target="_blank" rel="noopener">Hugging Face — bitsandbytes ↗</a></p>
327
+ <p><a href="https://huggingface.co/docs/transformers/quantization/gptq" target="_blank" rel="noopener">Hugging Face — GPTQ ↗</a></p>
328
+ <p><a href="https://huggingface.co/docs/transformers/quantization/awq" target="_blank" rel="noopener">Hugging Face — AWQ ↗</a></p>
329
+ <p><a href="https://github.com/ggml-org/llama.cpp/tree/master/tools/quantize" target="_blank" rel="noopener">llama.cpp — quantization tools ↗</a></p>
330
+ </section>
331
+
332
+ <footer>
333
+ <strong style="color:white">Open Weight</strong><br>
334
+ Open weights. Portable models. Deployable AI.<br><br>
335
+ Collaboration: open-weight AI, model infrastructure, inference, deployment, research and ecosystem partnerships.<br>
336
+ Contact: <a href="mailto:agenten@magenta.de">agenten@magenta.de</a>
337
+ </footer>
338
+ </div>
339
+
340
+ <script>
341
+ function updateMemory(){
342
+ const p = Math.max(0, parseFloat(document.getElementById('params').value)||0);
343
+ const b = Math.max(0, parseFloat(document.getElementById('bits').value)||0);
344
+ const gb = p * b / 8;
345
+ document.getElementById('memoryOut').textContent = gb.toFixed(2) + ' GB';
346
+ }
347
+ document.getElementById('params').addEventListener('input',updateMemory);
348
+ document.getElementById('bits').addEventListener('change',updateMemory);
349
+
350
+ const answers={
351
+ hf:`<strong>Start by evaluating bitsandbytes.</strong><p class="muted">Its 4-bit and 8-bit Transformers integration makes it a practical entry point when your model and hardware are supported.</p>`,
352
+ local:`<strong>Investigate GGUF / llama.cpp quantization.</strong><p class="muted">The ecosystem offers many low-bit formats for local inference across supported CPUs, GPUs and Apple Silicon workflows.</p>`,
353
+ gpu:`<strong>Choose the serving runtime first.</strong><p class="muted">Then compare the methods it accelerates well — such as supported FP8, GPTQ, AWQ, bitsandbytes or other native quantization paths.</p>`,
354
+ tune:`<strong>Look closely at 4-bit PEFT / QLoRA workflows.</strong><p class="muted">bitsandbytes is a common entry point because it combines low-bit loading with trainable adapter parameters.</p>`
355
+ };
356
+ document.querySelectorAll('.tab').forEach(btn=>{
357
+ btn.addEventListener('click',()=>{
358
+ document.querySelectorAll('.tab').forEach(x=>x.classList.remove('active'));
359
+ btn.classList.add('active');
360
+ document.getElementById('answer').innerHTML=answers[btn.dataset.answer];
361
+ });
362
+ });
363
+ updateMemory();
364
+ </script>
365
+ </body>
366
+ </html>