antjeworring commited on
Commit
3f803c6
·
verified ·
1 Parent(s): 11fa19c

Publish Zen6 comprehensive model card and specifications

Browse files
Files changed (1) hide show
  1. README.md +139 -0
README.md ADDED
@@ -0,0 +1,139 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: qwen-community-1.0
4
+ language:
5
+ - en
6
+ - zh
7
+ pipeline_tag: text-generation
8
+ tags:
9
+ - zen6
10
+ - zen6-coder
11
+ - agentic-coding
12
+ - moe
13
+ - qwen3.8-flash-next
14
+ - unsloth
15
+ - halogen
16
+ - strix-halo
17
+ - mtp
18
+ base_model:
19
+ - Qwen/Qwen3.8-Flash-Next
20
+ - unsloth/Qwen3.8-Flash-Next-GGUF
21
+ ---
22
+
23
+ <div align="center">
24
+
25
+ # Zen6 Coder: 180B Frontier Agentic MoE
26
+
27
+ **125B Base (6B Active) | 51B N-Gram Embedding | 4B MTP Drafter | 62.5 SWE-bench Pro**
28
+
29
+ [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-zenlm%2Fzen6--coder-blue)](https://huggingface.co/zenlm/zen6-coder)
30
+ [![GitHub](https://img.shields.io/badge/GitHub-zenlm%2Fzen6--coder-black)](https://github.com/zenlm/zen6-coder)
31
+
32
+ </div>
33
+
34
+ ---
35
+
36
+ ## Architectural Highlights
37
+
38
+ **Zen6 Coder** is built on the next-generation **Qwen3.8-Flash-Next** architecture, representing a fundamental redesign of modern agentic language models:
39
+
40
+ - **180 Billion Total Parameters**:
41
+ - **125B Base Language Model** with only **6B Activated Parameters** per token (10 routed experts + 1 shared expert out of 512 total experts).
42
+ - **51B N-Gram Embedding Table** (20,000,000 bigrams/trigrams injected at layer 2) enabling ultra-dense lexical memory without compute overhead.
43
+ - **4B Multi-Token Prediction (MTP) Head** (1 dedicated layer trained with multi-step prediction) delivering 1.3x–1.7x speculative acceleration out of the box.
44
+ - **Hybrid Attention with QSA (Qwen Sparse Attention)**:
45
+ - 48 Layers arranged as $12 \times [3 \times (\text{Gated DeltaNet} \to \text{MoE}) \to 1 \times (\text{QSA} \to \text{MoE})]$.
46
+ - **Gated DeltaNet**: 48 linear attention heads for V, 16 heads for QK (head dim 128) handling constant-memory linear sequence progression.
47
+ - **QSA**: 24 Query heads, 2 KV heads (head dim 256, RoPE dim 64) with an MQA Indexer (4 Query / 1 Shared Key, budget 512 micro-blocks / 2048 tokens).
48
+ - **Gated Residuals**: 4 residual branches modulated by data-dependent read and write gates with bottleneck rank 320.
49
+ - **Context Length**: 262,144 tokens native, extensible to 1,000,000 tokens via YaRN (`rope_theta: 10000000, factor: 4.0`).
50
+
51
+ ---
52
+
53
+ ## State-of-the-Art Coding & Agent Benchmarks
54
+
55
+ Zen6 Coder establishes new state-of-the-art benchmarks in real-world software engineering and agentic coding:
56
+
57
+ | Benchmark | Zen6 Coder (Qwen3.8-Flash-Next) | Claude-Opus-4.6 (Max) | DeepSeek-V4-Flash-0731 | Qwen3.8-27B |
58
+ | :--- | :---: | :---: | :---: | :---: |
59
+ | **SWE-bench Pro** | **62.5%** | 53.4% | 56.0% | 61.7% |
60
+ | **DeepSWE 1.1** | **58.7%** | — | 54.4% | 42.2% |
61
+ | **SWE-bench Multilingual** | **81.0%** | 77.5% | — | 73.8% |
62
+ | **LiveCodeBench v6** | **91.9%** | 88.8% | 90.6% | 90.3% |
63
+ | **NL2Repo-Bench** | **48.1%** | 47.6% | 54.2% | 42.3% |
64
+ | **GPQA Diamond** | **91.7%** | 91.3% | 90.8% | 89.2% |
65
+ | **Toolathlon Verified (Pass@1)** | **73.5%** | — | 70.3% | 67.1% |
66
+ | **CoWorkBench** | **73.9%** | 68.2% | 45.1% | 70.7% |
67
+
68
+ ---
69
+
70
+ ## Model Weights & Formats
71
+
72
+ This repository distributes Zen6 Coder in two primary formats:
73
+
74
+ ### 1. Unsloth Dynamic GGUF (`UD-IQ4_XS`) + MTP
75
+ - **`UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf`** (10.9 MB)
76
+ - **`UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00002-of-00003.gguf`** (49.8 GB)
77
+ - **`UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00003-of-00003.gguf`** (43.8 GB)
78
+ - **`MTP/mtp-Qwen3.8-Flash-Next-Q8_0.gguf`** (Dedicated 4B MTP draft head)
79
+
80
+ ### 2. Halogen W4B Format (AMD Strix Halo Native)
81
+ Optimized for AMD Ryzen AI Max+ 395 / Radeon 8060S (gfx1151) with ROCm and Halogen resumable prompt-state caching.
82
+
83
+ ---
84
+
85
+ ## Hardware Benchmarks
86
+
87
+ ### AMD Strix Halo (8060S / 128GB Unified Memory)
88
+ | Context Length | Cold Prefill | Halogen Warm Resume | Speedup |
89
+ | :---: | :---: | :---: | :---: |
90
+ | **512 tokens** | 454.4 tok/s | **0.1 ms** | **7.89x** |
91
+ | **2,048 tokens** | 959.1 tok/s | **0.1 ms** | **14.08x** |
92
+ | **8,192 tokens** | 1,298.4 tok/s | **0.1 ms** | **36.31x** |
93
+ | **16,384 tokens** | 1,373.4 tok/s | **0.1 ms** | **60.47x** |
94
+ | **32,768 tokens** | 1,451.8 tok/s | **0.1 ms** | **79.45x** |
95
+
96
+ ---
97
+
98
+ ## Serving Instructions
99
+
100
+ ### Option A: AMD Strix Halo (Halogen Engine)
101
+ ```bash
102
+ sudo podman run -d --name halogen --device=/dev/kfd --device=/dev/dri \
103
+ -v /models:/models -p 8731:8731 halogen:latest \
104
+ --model /models/qwen38-flash-next-w4b.hgn \
105
+ --port 8731 --max-tokens-cap 65536
106
+ ```
107
+
108
+ ### Option B: Cross-Platform Llama.cpp with MTP Speculative Decoding
109
+ ```bash
110
+ llama-server \
111
+ -m UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf \
112
+ --draft-model MTP/mtp-Qwen3.8-Flash-Next-Q8_0.gguf \
113
+ --draft-max 3 \
114
+ -c 262144 \
115
+ --port 8000
116
+ ```
117
+
118
+ ### Option C: Pure-Rust `hanzo-engine`
119
+ ```bash
120
+ hanzo-engine serve \
121
+ --model zenlm/zen6-coder \
122
+ --format gguf \
123
+ --mtp MTP/mtp-Qwen3.8-Flash-Next-Q8_0.gguf \
124
+ --context-window 262144 \
125
+ --port 8000
126
+ ```
127
+
128
+ ---
129
+
130
+ ## Citation
131
+
132
+ ```bibtex
133
+ @techreport{zenlm2026zen6coder,
134
+ title={Zen6 Coder: 180B-Class Hybrid Gated DeltaNet Sparse Attention MoE for Frontier Agentic Software Engineering},
135
+ author={Hanzo AI and Zen LM Team},
136
+ year={2026},
137
+ publisher={Zen LM / Hanzo AI}
138
+ }
139
+ ```