1337Hero commited on
Commit
b3ac135
Β·
verified Β·
1 Parent(s): 6ec1471

Add wikitext-2 perplexity vs BF16 source; revise file recommendation

Browse files
Files changed (1) hide show
  1. README.md +49 -17
README.md CHANGED
@@ -37,21 +37,29 @@ a 34.66B-parameter MoE coding model (256 experts, 8 active, 256K context,
37
  > anyway, do not trust the output.
38
 
39
  > [!WARNING]
40
- > Both files load, generate coherent output, and were throughput-benchmarked
41
- > on RDNA4 `gfx1201`. That is the **only** validation performed. No Strix Halo
42
- > testing and **no quality evaluation of any kind** β€” see [What was not
 
43
  > measured](#what-was-not-measured) before relying on either file.
44
 
45
  ## Which file?
46
 
47
- | File | Size | Effective BPW | Pick it if |
48
- | --- | ---: | ---: | --- |
49
- | `KAT-Coder-V2.5-Dev-Q4_0_ROCMFP4_STRIX_LEAN.gguf` | 17.32 GiB | 4.29 | You want the smaller, faster file. Recommended default. |
50
- | `KAT-Coder-V2.5-Dev-Q4_0_ROCMFP4.gguf` | 21.18 GiB | 5.25 | You want more bits on the expert down-projections. |
51
 
52
- `STRIX_LEAN` is 18% smaller *and* measurably faster on both backends, so it is
53
- the recommended starting point despite the Strix-oriented name. The recipe was
54
- tuned on `gfx1151`; nothing about the file format is Strix-specific.
 
 
 
 
 
 
 
55
 
56
  ## Why the sizes differ from the nominal BPW
57
 
@@ -86,17 +94,41 @@ Two results worth acting on:
86
  - **Use Vulkan on this hardware.** Vulkan decodes roughly **2Γ— faster** than
87
  HIP/ROCm for both files (122 vs 59 t/s on `STRIX_LEAN`) and also leads on
88
  prompt fill. This matches ROCmFPX's own Strix Halo findings.
89
- - **`STRIX_LEAN` wins on both axes.** It is 18% smaller *and* faster β€”
90
- +13% decode and +5% prefill on Vulkan, +13% decode and +45% prefill on ROCm.
 
91
 
92
  No control quant (Q4_K_M or similar) was benchmarked, so these numbers compare
93
  the two ROCmFP4 files against each other, not against ordinary GGUF quants.
94
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
95
  ## What was not measured
96
 
97
- - **Output quality.** No perplexity, KL-divergence, HumanEval, or MBPP
98
- comparison against the BF16 source. Neither file has been quality-checked at
99
- all.
 
100
  - **Agentic and tool-calling behavior**, which is the point of a coding model.
101
  Untested.
102
  - **Any hardware other than `gfx1201`.** Not tested on Strix Halo, RDNA3,
@@ -123,7 +155,7 @@ automatically. `gfx1200` builds are **not** interchangeable on these cards.
123
 
124
  ```bash
125
  ./build-rdna4/bin/llama-server \
126
- -m KAT-Coder-V2.5-Dev-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
127
  -dev Vulkan0 \
128
  -ngl 999 \
129
  -fa on \
@@ -191,7 +223,7 @@ and the converted GGUFs contain no multimodal projector. Despite the
191
 
192
  - Requires the ROCmFPX fork; no upstream llama.cpp compatibility.
193
  - Validated on exactly one `gfx1201` host, batch 1, shallow context.
194
- - No quality evaluation of any kind has been published for these artifacts.
195
  - 34.66B MoE: needs ~18–22 GB for weights plus KV cache. Comfortable on a
196
  32 GB card, tight on 24 GB with meaningful context.
197
 
 
37
  > anyway, do not trust the output.
38
 
39
  > [!WARNING]
40
+ > Validation was performed on RDNA4 `gfx1201` only: both files load, generate
41
+ > coherent output, were throughput-benchmarked, and were measured against the
42
+ > BF16 source for wikitext-2 perplexity. No Strix Halo testing and **no
43
+ > code-specific or agentic evaluation** β€” see [What was not
44
  > measured](#what-was-not-measured) before relying on either file.
45
 
46
  ## Which file?
47
 
48
+ | File | Size | Effective BPW | Wikitext-2 PPL | Pick it if |
49
+ | --- | ---: | ---: | ---: | --- |
50
+ | `KAT-Coder-V2.5-Dev-Q4_0_ROCMFP4.gguf` | 21.18 GiB | 5.25 | 6.9182 (+1.38%) | You care about output quality. **Recommended for coding.** |
51
+ | `KAT-Coder-V2.5-Dev-Q4_0_ROCMFP4_STRIX_LEAN.gguf` | 17.32 GiB | 4.29 | 7.1079 (+4.16%) | You need the smaller file or the extra decode speed. |
52
 
53
+ This is a **real tradeoff, not a clean win for either file.** `STRIX_LEAN` is
54
+ 18% smaller and 13% faster at decode, but gives up three times as much
55
+ perplexity against the BF16 source. For a coding model β€” where a single wrong
56
+ token breaks a program β€” the plain `Q4_0_ROCMFP4` is the safer default, and
57
+ 21.18 GiB still fits a 32 GB card comfortably.
58
+
59
+ Take `STRIX_LEAN` if you are memory-constrained (24 GB cards), or if you are
60
+ throughput-bound and have validated that the quality holds on your own tasks.
61
+ Its recipe was tuned on `gfx1151`; nothing about the file format is
62
+ Strix-specific.
63
 
64
  ## Why the sizes differ from the nominal BPW
65
 
 
94
  - **Use Vulkan on this hardware.** Vulkan decodes roughly **2Γ— faster** than
95
  HIP/ROCm for both files (122 vs 59 t/s on `STRIX_LEAN`) and also leads on
96
  prompt fill. This matches ROCmFPX's own Strix Halo findings.
97
+ - **`STRIX_LEAN` is the faster file** β€” +13% decode and +5% prefill on Vulkan,
98
+ +13% decode and +45% prefill on ROCm β€” but see the quality section below
99
+ before choosing it on speed alone.
100
 
101
  No control quant (Q4_K_M or similar) was benchmarked, so these numbers compare
102
  the two ROCmFP4 files against each other, not against ordinary GGUF quants.
103
 
104
+ ## Measured quality β€” wikitext-2 perplexity
105
+
106
+ `llama-perplexity`, full wikitext-2 test set (580 chunks), `-c 512 -b 512`,
107
+ FlashAttention on, Vulkan. The BF16 source GGUF was measured on the same host
108
+ with the same settings, split across three GPUs.
109
+
110
+ | File | BPW | PPL | Ξ” vs BF16 |
111
+ | --- | ---: | ---: | ---: |
112
+ | `KAT-Coder-V2.5-Dev-BF16.gguf` (source) | 16.01 | 6.8237 Β± 0.04537 | β€” |
113
+ | `Q4_0_ROCMFP4` | 5.25 | 6.9182 Β± 0.04607 | **+1.38%** |
114
+ | `Q4_0_ROCMFP4_STRIX_LEAN` | 4.29 | 7.1079 Β± 0.04762 | **+4.16%** |
115
+
116
+ Both quants land where you would expect for their bit budgets, and neither is
117
+ degenerate. The gap between them is larger than the error bars, so it is a
118
+ real difference and not measurement noise: `STRIX_LEAN` buys its 18% size
119
+ reduction with roughly 3Γ— the perplexity cost.
120
+
121
+ Perplexity is a weak proxy for coding ability. It measures next-token
122
+ prediction on English Wikipedia, not code correctness or tool-call formatting.
123
+ Treat it as a floor check β€” it rules out a broken quantization, it does not
124
+ establish that either file codes as well as the source.
125
+
126
  ## What was not measured
127
 
128
+ - **Coding ability.** No HumanEval, MBPP, or any code benchmark. Wikitext-2
129
+ perplexity was measured (see above), but it does not measure code
130
+ correctness.
131
+ - **KL-divergence** against the BF16 source. Perplexity only.
132
  - **Agentic and tool-calling behavior**, which is the point of a coding model.
133
  Untested.
134
  - **Any hardware other than `gfx1201`.** Not tested on Strix Halo, RDNA3,
 
155
 
156
  ```bash
157
  ./build-rdna4/bin/llama-server \
158
+ -m KAT-Coder-V2.5-Dev-Q4_0_ROCMFP4.gguf \
159
  -dev Vulkan0 \
160
  -ngl 999 \
161
  -fa on \
 
223
 
224
  - Requires the ROCmFPX fork; no upstream llama.cpp compatibility.
225
  - Validated on exactly one `gfx1201` host, batch 1, shallow context.
226
+ - Quality evidence is wikitext-2 perplexity only; no code or agentic evals.
227
  - 34.66B MoE: needs ~18–22 GB for weights plus KV cache. Comfortable on a
228
  32 GB card, tight on 24 GB with meaningful context.
229