ASHQ1-Remix: Attribution-Calibrated Hybrid Quantization Suite (v2.4.2)

Engine Lineage

ASHQ1-Remix is an empirical, activation-aware quantization framework for GGUF language models and speculative decoding modules.

Standard uniform quantization applies uniform bitwidths or static heuristics across model layers. ASHQ1-Remix analyzes real activation distributions via llama-imatrix, calculates empirical sensitivity curves ($\Delta\text{KLD}/\text{MiB}$) across distinct transformer and hybrid components (attention projections, recurrent state-space channels, SwiGLU MLPs, embedding matrices), and allocates optimal bitwidths under disciplined footprint targets.


πŸͺœ The Mod-3 Arithmetic Ladder

ASHQ1 tiers adhere to the standardized Mod-3 Ladder (% of unquantized BF16 source):

Tier Ratio Allocation Strategy Primary Operational Purpose
Pico 24% Dissolution-floor protection (IQ3_M / IQ3_S, scarce attention Q5_K) Budget edge service tier (ASHQ1_INCLUDE_PICO=1).
Nano 27% Low-band knapsack (IQ3_S–IQ4_XS base, tied readouts Q6_K) Compact edge deployment tier below stock IQ4_XS.
Mini 30% Allocator knapsack β€” or L4 flat on GDN hybrids (IQ4_XS + recurrent floors + full-attn band) Recommended balanced entry tier for models $\ge$ 9B.
Compact 33% Balanced knapsack β€” or flat Q4_K_M + full-attn band (+ recurrent floors on GDN hybrids) Recommended daily driver for 3B–4B architectures.
Quality 36% Structural flat Q5_K_M + imatrix (Law L6) General production tier with near-lossless fidelity and optimized inference speed.
Precision 42% Flat Q6_K base + readout Q8_0 lever (tied: token_embd Β· untied: lm_head β€” Law L12) Coding, mathematics, and high-entropy reasoning tasks.
Fidelity 48% Q6_K baseline + Q8_0 imatrix pockets & readouts High-fidelity archival tier (ASHQ1_INCLUDE_FIDELITY=1).

πŸ“Š Empirical Benchmarks

Evaluations run on wiki.test.raw (ctx=512, 64 chunks, span = 32,768 tokens) using Flash-Attention against unquantized BF16 reference logits (kld-bf16.dat).

Evaluation Legend

  • ⭐ Recommended: Best balance of generation quality and memory efficiency for targeted deployment hardware (e.g. Mini for 9B models on 8 GB VRAM; Quality for small models).
  • πŸ₯ˆ Second Choice: Alternative tier offering either higher fidelity or extra KV cache capacity.
  • βœ— Not Recommended: Tiers displaying quality degradation ($\text{KLD} > 0.20$, $\text{top-p} < 80%$) or underperforming standard baselines.
  • Classic Baselines: Standard (i)QN_X_Y quants appear for comparative evaluation.
  • Constructibility Floors: Omitted ASHQ1 tiers in specific tables indicate structural floor constraints where compliant generation cannot be formed within ladder budgets.
  • Over-confidence Artifacts (†): Nominal PPL falling below the unquantized base denotes probability distribution compression; ranking relies on KLD and top-p agreement.

1. Ornith-1.5-9B (GDN Hybrid, Untied 248k Vocab)

BF16 reference: 17,555 MiB Β· Recommended target: 8+ GB VRAM hardware.

Model Size PPL KLD RMS Ξ”p top-p Speed
Q8_0 (stock) 9333 MiB 9.0824 0.0058 2.26% 97.9% 244 t/s
Fidelity-48pc 8436 MiB 9.0227† 0.0087 2.58% 97.2% 260 t/s
Precision-42pc 7554 MiB 8.9957† 0.0094 2.72% 96.8% 324 t/s
Q6_K-imx (stock) 7209 MiB 8.9836† 0.0107 2.90% 96.4% 320 t/s
Quality-36pc 6388 MiB 8.9895† 0.0241 4.22% 94.5% 1313 t/s
Q5_K_M-imx (stock) 6335 MiB 8.6345† 0.0665 6.48% 90.7% 1219 t/s
Compact-33pc ⭐ 5945 MiB 9.1332 0.0319 4.88% 93.1% 1219 t/s
Mini-30pc πŸ₯ˆ 5566 MiB 9.0189† 0.0425 5.65% 91.8% 1384 t/s
IQ4_XS-imx (stock) 5080 MiB 9.2758 0.0538 6.29% 90.7% 1506 t/s
Nano-27pc 4750 MiB 9.5968 0.0802 7.57% 87.9% 1249 t/s
Pico-24pc 4389 MiB 9.5871 0.1202 9.25% 85.1% 1384 t/s
IQ3_M-imx (stock) 4313 MiB 9.5973 0.1421 10.46% 84.3% 1384 t/s

2. Qwen3.8-4B-Distill (Linear Attention / GDN Hybrid)

BF16 reference: 8034 MiB Β· Compact omitted due to floor boundaries Β· Recommended target: 4+ GB VRAM hardware.

Model Size PPL KLD RMS Ξ”p top-p Speed
Q8_0 (stock) 4275 MiB 8.6208 0.0008 0.85% 98.5% 1969 t/s
Fidelity-48pc 3866 MiB 8.6382 0.0015 1.15% 98.1% 1347 t/s
Precision-42pc 3519 MiB 8.6473 0.0020 1.32% 97.8% 1707 t/s
Q6_K-imx (stock) 3304 MiB 8.6572 0.0026 1.47% 97.2% 1600 t/s
Quality-36pc ⭐ 2966 MiB 8.6863 0.0063 2.18% 96.0% 1829 t/s
Q5_K_M-imx (stock) 2933 MiB 8.6820 0.0081 2.45% 95.3% 1829 t/s
Mini-30pc πŸ₯ˆ 2585 MiB 8.7532 0.0189 3.74% 93.0% 1829 t/s
IQ4_XS-imx (stock) 2398 MiB 8.8026 0.0262 4.40% 92.0% 1896 t/s
Nano-27pc 2245 MiB 9.1127 0.0527 6.61% 89.1% 1707 t/s
Pico-24pc 2168 MiB 9.1831 0.0631 7.12% 88.1% 1829 t/s
IQ3_M-imx (stock) 2063 MiB 9.3394 0.0799 8.27% 86.6% 1652 t/s

3. Nanbeige4.2-3B (Looped Dense Trunk, Untied 166k Vocab)

BF16 reference: 7957 MiB Β· Readout accounts for 24.5% of total mass Β· Recommended target: 6+ GB VRAM hardware.

Model Size PPL KLD RMS Ξ”p top-p Speed
Q8_0 (stock) 4229 MiB 34.4742 0.0105 2.39% 95.6% 353 t/s
Fidelity-48pc πŸ₯ˆ 3824 MiB 34.3246† 0.0247 3.42% 94.0% 371 t/s
Precision-42pc ⭐ 3384 MiB 34.3791† 0.0331 3.94% 92.5% 382 t/s
Q6_K-imx (stock) 3266 MiB 34.4411† 0.0335 3.99% 92.4% 406 t/s
Quality-36pc 2849 MiB 34.6224 0.0743 5.80% 88.3% 466 t/s
Q5_K_M-imx (stock) 2849 MiB 34.6224 0.0743 5.80% 88.3% 470 t/s
Compact-33pc 2630 MiB 34.6636 0.1441 7.98% 83.8% 512 t/s
Mini-30pc 2391 MiB 33.5214† 0.1728 8.66% 81.8% 557 t/s
IQ4_XS-imx (stock) 2268 MiB 34.8506 0.1813 9.17% 81.2% 457 t/s
Nano-27pc βœ— 2152 MiB 33.7551† 0.2191 10.12% 78.6% 453 t/s
IQ3_M-imx (stock) 1985 MiB 36.4366 0.4492 14.31% 70.5% 539 t/s
Pico-24pc βœ— 1913 MiB 37.8642 0.3878 13.31% 72.6% 545 t/s

4. Spark-X2.5-4B (Dense SWA-Hybrid 3:1, Tied 131k Vocab)

BF16 reference: 7849 MiB Β· Recommended target: 6+ GB VRAM hardware.

Model Size PPL KLD RMS Ξ”p top-p Speed
Q8_0 (stock) 4172 MiB 34.4529 0.0075 2.03% 96.4% 2844 t/s
Fidelity-48pc πŸ₯ˆ 3772 MiB 33.8560† 0.0185 2.85% 94.0% 2133 t/s
Precision-42pc ⭐ 3359 MiB 33.2561† 0.0213 3.30% 93.5% 1766 t/s
Q6_K-imx (stock) 3223 MiB 33.2700† 0.0269 3.73% 92.5% 2327 t/s
Quality-36pc 2840 MiB 34.9522 0.0719 5.88% 88.2% 2327 t/s
Q5_K_M-imx (stock) 2840 MiB 34.9522 0.0719 5.88% 88.2% 2438 t/s
Compact-33pc 2593 MiB 38.5817 0.1688 8.67% 82.0% 1829 t/s
Mini-30pc 2343 MiB 38.9367 0.2012 9.85% 79.9% 2226 t/s
IQ4_XS-imx (stock) 2266 MiB 41.9845 0.2586 10.90% 77.5% 2560 t/s
Nano-27pc βœ— 2124 MiB 40.6742 0.3281 12.32% 75.3% 2438 t/s
Pico-24pc βœ— 1982 MiB 31.7680† 0.3910 13.52% 72.4% 1766 t/s
IQ3_M-imx (stock) 1949 MiB 41.0486 0.4913 15.49% 69.1% 2327 t/s

5. TwIL-LM3 (Dense SmolLM3-Arch, Tied 128k Vocab)

BF16 reference: 5873 MiB Β· All ASHQ1 rungs strictly monotone.

Model Size PPL KLD RMS Ξ”p top-p Speed
Q8_0 (stock) 3124 MiB 9.8138 0.0017 1.06% 97.4% 2560 t/s
Fidelity-48pc 2826 MiB 9.8173 0.0024 1.23% 97.0% 2560 t/s
Precision-42pc 2474 MiB 9.8206 0.0037 1.50% 96.4% 2226 t/s
Q6_K-imx (stock) 2414 MiB 9.8629 0.0088 2.54% 93.9% 2695 t/s
Quality-36pc πŸ₯ˆ 2111 MiB 9.9407 0.0144 3.20% 92.5% 3012 t/s
Q5_K_M-imx (stock) 2111 MiB 9.9407 0.0144 3.20% 92.5% 3012 t/s
Compact-33pc ⭐ 1945 MiB 9.9726 0.0242 3.93% 91.2% 2560 t/s
Mini-30pc 1769 MiB 10.0104 0.0327 4.56% 89.9% 2695 t/s
IQ4_XS-imx (stock) 1644 MiB 10.1153 0.0381 4.92% 89.4% 2844 t/s
Nano-27pc 1593 MiB 10.2711 0.0500 5.71% 88.0% 2844 t/s
Pico-24pc 1445 MiB 10.7211 0.0974 7.78% 84.7% 2695 t/s
IQ3_M-imx (stock) 1401 MiB 11.0072 0.1064 8.69% 83.4% 2695 t/s

6. LFM2.5-2.6B (Shortconv-Mixer Hybrid, Tied Readout)

BF16 reference: 5153 MiB Β· Usable floor sits at Mini-30.

Model Size PPL KLD RMS Ξ”p top-p Speed
Q8_0 (stock) 2742 MiB 55.0069 0.0030 1.22% 97.4% 3413 t/s
Fidelity-48pc 2481 MiB 55.6994 0.0074 1.97% 96.0% 2844 t/s
Precision-42pc πŸ₯ˆ 2199 MiB 55.2931 0.0087 2.20% 95.7% 2695 t/s
Q6_K-imx (stock) 2119 MiB 55.6407 0.0122 2.54% 94.7% 2844 t/s
Quality-36pc ⭐ 1850 MiB 54.9504† 0.0347 3.95% 91.1% 3200 t/s
Q5_K_M-imx (stock) 1850 MiB 54.9504† 0.0347 3.95% 91.1% 2844 t/s
Compact-33pc 1708 MiB 53.0954† 0.0734 6.01% 87.5% 2844 t/s
Mini-30pc 1554 MiB 49.8535† 0.1291 8.10% 83.4% 2844 t/s
IQ4_XS-imx (stock) 1447 MiB 55.5126 0.1486 8.59% 81.8% 2560 t/s
Nano-27pc βœ— 1399 MiB 58.9388 0.2130 10.16% 78.3% 3012 t/s
Pico-24pc βœ— 1276 MiB 54.4368† 0.3644 12.94% 72.7% 3200 t/s
IQ3_M-imx (stock) 1225 MiB 59.3136 0.3884 13.44% 72.2% 3012 t/s

7. MiniCPM5-2B (Dense llama-arch, Untied 130k Readout)

BF16 reference: 4806 MiB Β· Newly benchmarked in release 2.4.2.

Model Size PPL KLD RMS Ξ”p top-p Speed
Q8_0 (stock) 2556 MiB 15.0634 0.0016 1.01% 97.7% 3939 t/s
Fidelity-48pc 2312 MiB 15.0764 0.0031 1.41% 96.7% 3657 t/s
Precision-42pc πŸ₯ˆ 2036 MiB 15.0718 0.0057 1.91% 95.9% 3657 t/s
Q6_K-imx (stock) 1974 MiB 15.0642 0.0061 1.98% 95.7% 3657 t/s
Quality-36pc ⭐ 1724 MiB 15.2230 0.0189 3.43% 92.6% 3200 t/s
Q5_K_M-imx (stock) 1724 MiB 15.2230 0.0189 3.43% 92.6% 3657 t/s
Compact-33pc 1591 MiB 15.5574 0.0533 5.63% 88.5% 3012 t/s
Mini-30pc βœ— 1447 MiB 15.8384 0.0772 6.79% 85.7% 3657 t/s
IQ4_XS-imx (stock) 1358 MiB 15.7752 0.0757 6.76% 86.0% 3413 t/s
Nano-27pc 1303 MiB 15.9764 0.0959 7.63% 83.8% 3939 t/s
IQ3_M-imx (stock) 1170 MiB 17.7124 0.2158 11.85% 77.1% 3200 t/s
Pico-24pc βœ— 1158 MiB 17.8654 0.2180 11.39% 77.0% 3413 t/s

8. ReaderLM-v2 (Dense HTML/Markdown Extractor, 1.5B, 75% FFN Mass)

BF16 reference: 2950 MiB Β· Law L2 validated (aligned Compact floor).

Model Size PPL KLD RMS Ξ”p top-p Speed
Q8_0 (stock) 1570 MiB 15.9804 0.0022 1.15% 97.6% 4655 t/s
Fidelity-48pc 1422 MiB 15.9874 0.0044 1.58% 96.4% 4267 t/s
Precision-42pc πŸ₯ˆ 1268 MiB 15.9426† 0.0067 1.98% 95.6% 3939 t/s
Q6_K-imx (stock) 1214 MiB 15.9578† 0.0074 2.08% 95.2% 4267 t/s
Quality-36pc ⭐ 1073 MiB 15.9779† 0.0223 3.71% 92.2% 4267 t/s
Q5_K_M-imx (stock) 1073 MiB 15.9779† 0.0223 3.71% 92.2% 4267 t/s
Compact-33pc 980 MiB 15.9972 0.0504 5.53% 88.6% 3939 t/s
Mini-30pc 891 MiB 16.0512 0.0751 6.61% 86.3% 4267 t/s
IQ4_XS-imx (stock) 854 MiB 16.0340 0.0815 6.96% 86.0% 4267 t/s
Nano-27pc 803 MiB 16.4019 0.1434 9.29% 81.5% 3939 t/s
Pico-24pc βœ— 763 MiB 17.0117 0.2145 11.49% 77.5% 3939 t/s
IQ3_M-imx (stock) 741 MiB 17.2831 0.2318 11.77% 77.1% 3939 t/s

9. MiniCPM5-1B (Dense llama-arch, Untied 130k Readout)

BF16 reference: 2066 MiB Β· Readout pair = 37% of total mass.

Model Size PPL KLD RMS Ξ”p top-p Speed
Q8_0 (stock) 1100 MiB 26.9046 0.0019 0.96% 97.2% 5689 t/s
Fidelity-48pc 997 MiB 27.0943 0.0056 1.66% 95.3% 5689 t/s
Precision-42pc πŸ₯ˆ 897 MiB 27.1550 0.0074 1.84% 94.5% 5689 t/s
Q6_K-imx (stock) 851 MiB 27.1517 0.0079 1.91% 94.3% 5689 t/s
Quality-36pc ⭐ 750 MiB 27.6966 0.0279 3.70% 90.1% 5120 t/s
Q5_K_M-imx (stock) 750 MiB 27.6966 0.0279 3.70% 90.1% 5689 t/s
Compact-33pc 687 MiB 27.6945 0.0775 6.01% 83.7% 5120 t/s
Mini-30pc 625 MiB 28.8936 0.1154 7.27% 80.8% 5120 t/s
IQ4_XS-imx (stock) 609 MiB 29.1625 0.1096 7.13% 81.0% 5689 t/s
Nano-27pc 563 MiB 28.9311 0.1355 7.98% 78.9% 5689 t/s
IQ3_M-imx (stock) 536 MiB 36.0620 0.3241 13.66% 69.2% 5120 t/s
Pico-24pc βœ— 503 MiB 37.4572 0.3886 14.66% 66.2% 5689 t/s

10. OvisOCR2 (Vision-Aligned Qwen3.5-Hybrid Backbone, Tied 248k Vocab)

BF16 reference: 1446 MiB Β· Compact and Mini omitted due to floor constraints.

Model Size PPL KLD RMS Ξ”p top-p Speed
Q8_0 (stock) 774 MiB 28.7134 0.0009 0.61% 98.3% 3413 t/s
Fidelity-48pc 704 MiB 28.7754 0.0019 0.96% 97.3% 2560 t/s
Precision-42pc πŸ₯ˆ 670 MiB 28.8111 0.0023 1.04% 97.2% 3200 t/s
Q6_K-imx (stock) 601 MiB 28.8894 0.0033 1.25% 96.5% 3413 t/s
Quality-36pc ⭐ 556 MiB 29.0026 0.0077 1.93% 94.7% 3200 t/s
Q5_K_M-imx (stock) 551 MiB 29.0602 0.0088 2.02% 94.2% 3200 t/s
IQ4_XS-imx (stock) 481 MiB 30.5363 0.0305 3.94% 90.2% 3413 t/s
Nano-27pc 457 MiB 30.5109 0.0598 5.82% 87.2% 3200 t/s
Pico-24pc 442 MiB 31.9258 0.0699 6.16% 85.6% 3200 t/s
IQ3_M-imx (stock) 433 MiB 32.0551 0.0895 7.35% 84.2% 3200 t/s

11. PaddleOCR-VL-1.6 (Vision-Language OCR Decoder, ERNIE-4.5-0.3B)

BF16 reference: 892 MiB Β· Untied readout pair = 45% of total mass.

Model Size PPL KLD RMS Ξ”p top-p Speed
Q8_0 (stock) 475 MiB 544.1185 0.0178 1.67% 91.6% 7314 t/s
Fidelity-48pc πŸ₯ˆ 430 MiB 556.6220 0.0650 2.94% 87.0% 7314 t/s
Precision-42pc 392 MiB 522.0295† 0.0686 3.56% 83.4% 7314 t/s
Q6_K-imx (stock) 367 MiB 522.3200† 0.0694 3.57% 83.2% 7314 t/s
Quality-36pc 326 MiB 604.4922 0.1388 4.57% 76.6% 6400 t/s
Q5_K_M-imx (stock) 326 MiB 604.4922 0.1388 4.57% 76.6% 7314 t/s
IQ4_XS-imx (stock) 269 MiB 606.9490 0.5280 9.29% 61.8% 6400 t/s
IQ3_M-imx (stock) 239 MiB 768.5256 1.0750 14.12% 47.4% 6400 t/s

πŸ“‰ Sub-Nano Compendium (Unified 64-Chunk Protocol)

Nano→Pico transitions measured across all 11 model families:

Family Architecture Nano KLD Pico KLD Step (KLD) Pico Operational Verdict
OvisOCR2 Vision backbone 0.0598 0.0699 +16.9% βœ… Production Ready (442 MiB Β· beats IQ3_M 0.0895)
Spark-X2.5-4B SWA hybrid 0.3281 0.3910 +19.2% βœ— Degraded (Floor = Mini)
Qwen3.8-4B GDN hybrid 0.0527 0.0631 +19.7% βœ… Production Ready (Beats IQ3_M 0.0799)
ReaderLM-v2 Dense extractor 0.1434 0.2145 +49.6% βœ— Sub-80% top-p (Floor = Mini)
Ornith-1.5-9B GDN hybrid 0.0802 0.1202 +49.9% βœ… Fully Usable (Beats IQ3_M 0.1421)
LFM2.5-2.6B Shortconv mixer 0.2130 0.3644 +71.1% βœ— Exceeds 0.20 KLD (Floor = Mini)
Nanbeige4.2-3B Looped dense 0.2191 0.3878 +77.0% βœ— Exceeds 0.20 KLD (Floor = Mini)
PaddleOCR-VL-1.6 Dense untied OCR 0.6891 1.2883 +86.9% βœ— Severe collapse (Floor = Quality)
TwIL-LM3 Dense tied 0.0500 0.0974 +94.8% βœ… Maintains Fidelity (KLD < 0.10, top-p 84.7%)
MiniCPM5-2B Dense untied 0.0959 0.2180 +127.3% βœ— Exceeds 0.20 KLD (Floor = Nano)
MiniCPM5-1B Dense untied 0.1355 0.3886 +186.8% βœ— Collapse below constructive limit (Floor = Nano)

⚑ Speculative Decoding: DSpark2 Subsystem

ASHQ1-Remix integrates block-diffusion speculative draft distillation:

  1. DSpark2 Engine (02b_BF16-GGUF-DSpark2-creator.py): distills draft decoders with Markov and confidence heads on on-policy target rollouts, producing standalone GGUF modules.
  2. Acceptance Battery (03_dspark-acceptance-sweep.py): measures live generation throughput and speculative acceptance length under llama-server.

Empirical Result (Nanbeige): the ASHQ1-Compact draft provides +53% net throughput over trunk alone; proposal acceptance rate remains invariant to quantization (~`0.49` across BF16, Q8_0, and Compact).


⚑ Quick Start

# 1. Convert Hugging Face Safetensors to pristine BF16 GGUF
python 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py ./safetensors

# 2. Compute activation importance (default Bartowski v5 semantic dataset)
python 01_create-calibration-dataset-and-imatrix.py

# 3. Quantize full ASHQ1 ladder (Pico/Fidelity opt-in via env vars)
python 02_BF16-GGUF-to-ASHQ1.py

# 4. Benchmark output perplexity and KL-divergence
python 03_perplexity_test.py

πŸ“œ Citation & Credits

  • ASHQ1 (Autonomous Selective Hybrid Quantization) by wepiqx β€” priority-queue knapsack formulation, tied-group activation hashing, MSE scheduling.
  • Empero AI (Qwen3.8-27B-Ridge) β€” GDN state preservation (ssm_alpha/ssm_beta @ Q8_0) and native MTP draft heads.
  • Intel AutoRound β€” sign-gradient low-bit optimization with Hessian compensation.
  • llama.cpp by Georgi Gerganov & ggml contributors β€” GGUF/GGML runtime and tools.

License: apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Soulfate24/AutoRound-ASHQ1-Remix_Double-Quantization_Suite

Base model

wepiqx/ASHQ1
Finetuned
(2)
this model