Instructions to use maxwelhelp/llama.cpp-DFlash2-pascal6-optimized with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use maxwelhelp/llama.cpp-DFlash2-pascal6-optimized with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf maxwelhelp/llama.cpp-DFlash2-pascal6-optimized # Run inference directly in the terminal: llama cli -hf maxwelhelp/llama.cpp-DFlash2-pascal6-optimized
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf maxwelhelp/llama.cpp-DFlash2-pascal6-optimized # Run inference directly in the terminal: llama cli -hf maxwelhelp/llama.cpp-DFlash2-pascal6-optimized
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf maxwelhelp/llama.cpp-DFlash2-pascal6-optimized # Run inference directly in the terminal: ./llama-cli -hf maxwelhelp/llama.cpp-DFlash2-pascal6-optimized
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf maxwelhelp/llama.cpp-DFlash2-pascal6-optimized # Run inference directly in the terminal: ./build/bin/llama-cli -hf maxwelhelp/llama.cpp-DFlash2-pascal6-optimized
Use Docker
docker model run hf.co/maxwelhelp/llama.cpp-DFlash2-pascal6-optimized
- LM Studio
- Jan
- Ollama
How to use maxwelhelp/llama.cpp-DFlash2-pascal6-optimized with Ollama:
ollama run hf.co/maxwelhelp/llama.cpp-DFlash2-pascal6-optimized
- Unsloth Desktop
- Docker Model Runner
How to use maxwelhelp/llama.cpp-DFlash2-pascal6-optimized with Docker Model Runner:
docker model run hf.co/maxwelhelp/llama.cpp-DFlash2-pascal6-optimized
- Lemonade
How to use maxwelhelp/llama.cpp-DFlash2-pascal6-optimized with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull maxwelhelp/llama.cpp-DFlash2-pascal6-optimized
Run and chat with the model
lemonade run user.llama.cpp-DFlash2-pascal6-optimized-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
File size: 6,131 Bytes
81e58a0 7c97475 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 | ---
language:
- en
- ru
tags:
- dflash2
- speculative-decoding
- llama.cpp
- tesla-p40
- pascal
- gguf
license: apache-2.0
library_name: gguf
base_model: Qwen/Qwen3.8-27B-UD
---
# llama.cpp DFlash2 P40 (Pascal) optimization
Speculative decoding (DFlash2 block-diffusion drafter) tuned specifically for
**NVIDIA Tesla P40** (Pascal, compute capability 6.1, no FP16 tensor cores).
Fork of `ggml-org/llama.cpp` (PR #27342 lineage) with DFlash2 support, optimized
for the Pascal backend: int8 DP4A (not FP16), multi-stage quantization of the
draft, and an adaptive margin that shortens the verify batch where the
draft's selector is uncertain.
```
Mini-article: Title placeholder -> 30.26 tok/s on Pascal P40 with 27B dense
```
---
## What this repo is
A runnable, measured build of llama.cpp + DFlash2 speculative decoding on a
single Tesla P40 (24 GB, sm_61). Everything in this repo's `DFLASH2_RING_PORT.md`
is a real measurement on the P40; no other GPUs / cloud numbers are claimed.
Primary pairing tested:
- **Target:** Qwen3.8-27B-UD `Qwen3.8-27B-UD-Q4_K_XL.gguf` (qwen35 hybrid, Gated DeltaNet)
- **Draft:** DFlash2 `Qwen3.8-27B-DFlash2-q4mix-self.gguf` (our q4-mix quant)
---
## Best command (hit it)
See `run_best.sh`. Summary:
```bash
./build-p40-ring/bin/llama-server \
-m Qwen3.8-27B-UD-Q4_K_XL.gguf \
-md Qwen3.8-27B-DFlash2-q4mix-self.gguf \
-ngl 999 -ngld 999 -c 8192 -b 512 -ub 512 -np 1 \
--load-mode mlock --cache-ram 32768 --checkpoint-min-step 512 \
-fa 1 -ctk q8_0 -ctv q8_0 -ctkd q8_0 -ctvd q8_0 --kv-unified \
--spec-type draft-dflash \
--spec-draft-n-max 7 --spec-draft-n-min 1 --spec-draft-p-min 0.35 \
--spec-draft-ctx 0 --temp 0 --jinja --reasoning off -bs \
-lv 4 --host 0.0.0.0 --port 8080
```
Key flags:
- `--spec-type draft-dflash` - DFlash2 block-diffusion drafter
- `--spec-draft-n-max 7` - sweep optimum (Test 26: 7=30.26 > 8=30.18 > 5=27.4 > 4=27.8)
- `--spec-draft-p-min 0.35` - adaptive margin (top1-top2 on raw selector logits),
cuts the chain where the selector is a coin flip. Shortens the verify batch.
- `--spec-draft-n-min 1` - keep short chains alive
- `-bs` - backend (GPU) argmax sampling, no CPU logit readback
- `-fa 1`, `-ctk/ctv q8_0`, `--kv-unified` - flash attention + q8 KV
- `--reasoning off` - fastest mode; use `on` for real reasoning (see below)
To switch to the reasoning scenario replace `--reasoning off` with `--reasoning on`.
---
## Measured results (Tesla P40, single GPU)
All numbers are our own runs, logged in `DFLASH2_RING_PORT.md`.
### Code prompt, reasoning OFF, greedy
| Draft quant | size | tok/s | acceptance | mean len |
|---|---|---|---|---|
| **q4-mix (recommended)** | 1.2 GB | **30.26** | 0.83 | 5.25 |
| Q8 (lucebox) | 2.0 GB | ~30 | 0.85 | 5.27 |
| Q2-hybrid (fc Q2, head) | 0.96GB | 29.0 | 0.80 | 5.17 |
| Q2-all | 0.87GB | 27.1 | 0.76 | 4.69 |
Draft quant is q4-mix = backbone q4_0 + dflash.* heads q8_0 (protects the
selector). q4-mix wins: same acceptance as Q8 with a smaller, faster draft;
Q2 is slower because the 2-bit selector heads lose near-tie accuracy.
### n_max sweep (q4-mix, code, reasoning OFF)
| n_max | tok/s |
|---|---|
| 4 | 27.8 |
| 5 | 27.4 |
| **7** | **30.26** |
| 8 | 30.18 |
n_max=7 is the optimum; adaptive margin already caps the useful chain at ~5.25.
### Thinking scenario (code, reasoning ON, greedy)
- ours (DFlash2 q4-mix + margin): **~18 t/s** with real reasoning tokens.
- Luce chain with real thinking: 13.3 t/s (their spec is killed by a budget hook),
i.e. our port is **faster than Luce on the reasoning scenario**.
### Reference (same 27B, code, reasoning OFF)
- our DFlash2 q4-mix + margin: **30.26**; plain Luce chain: ~32; ngram-mod /
ngram-map-k4v on top of DFlash2: **28.2** (drags us down, not used).
---
## What was done / optimizations (P40 Pascal)
1. **MMQ DP4A (int8), no FP16.**
Pascal has no FP16 tensor cores. `ggml-cuda` selects the MMQ DP4A kernel
(`__dp4a`, cc >= 610) for quantized weights via
`(!fp16_mma_hardware_available(cc) || ...)`, compile arch `61`. Draft and
target both run the int8 matmul path. No FP16 path is used for quant weights.
2. **Draft quantization pipeline** (`server/scripts/quantize_dflash_draft.py`
+ `llama-quantize`):
- `q4-mix`: backbone q4_0, `dflash.*` heads + conv q8_0 -> 1.2 GB, acceptance-neutral.
- Q2 available only via `llama-quantize` (python-gguf has no K-quant *write*).
3. **DFlash2 adaptive chain (margin)** (`common/speculative.cpp`):
raw selector logits top1-top2 margin vs a threshold from `p_min`; stops the
chain when the pick is a coin flip. This is what shortens the verify batch
and makes n_max=7 optimal.
4. **DFlash2 is block-diffusion**: builds one noise block (id_last + n_max masks)
and decodes it in a **single pass** (`draft_dflash::draft`, one
`llama_decode`) - same approach as upstream PR #27342.
5. **Backend (GPU) argmax** (`-bs`): no CPU-side logit readback during draft/verify.
6. **Flash attention** (`-fa`) + `q8_0` KV for both models, unified KV.
7. **Duplicate draft quant tooling** and honest per-scenario comparison vs
Luce (see DFLASH2_RING_PORT.md).
---
## HF model
Our recommended draft (q4-mix, 1.2 GB):
- Filename: `Qwen3.8-27B-DFlash2-q4mix-self.gguf`
- HF repo: **https://huggingface.co/maxwelhelp/llama.cpp-DFlash2-pascal6-optimized**
Target: `Qwen3.8-27B-UD-Q4_K_XL.gguf` (Qwen3.8-27B-UD, qwen35 hybrid).
---
## Building for P40
```bash
cmake -B build-p40-ring -DCMAKE_CUDA_ARCHITECTURES=61 -DLLAMA_CUDA=ON -DGGML_CUDA=ON -DLLAMA_CURL=ON
cmake --build build-p40-ring --config Release -j
```
---
## Notes / honesty
- No cloud numbers; all measurements are local on the P40 (see the log).
- ngram-mod / ngram-map-k4v did **not** help on top of DFlash2 and are not used.
- On the reasoning scenario our port already beats Luce; q4-mix + margin is the
recommended release config.
---
License note: this is a private fork for experimental Pascal tuning; it wraps
upstream llama.cpp (MIT) and DFlash2 draft weights (Apache-2.0 / z-lab).
|