clef-flash-FP8-Dynamic
Unofficial FP8_DYNAMIC quantization of Cloudflare/clef-flash, a 9B decision model that scores every option of every typed question (choice / score / true-false) in one forward pass. This repo is not affiliated with or endorsed by Cloudflare.
It ships with a small vLLM plugin and a Jev/SystemOne-compatible HTTP server (POST /v1/systemone),
because stock vLLM cannot run Clef's custom joint-schema head. In short: near-lossless (98.8% top-1 agreement with BF16) and 1.5x BF16's two-GPU throughput on a single GPU.
Read before using
- Text and JSON state only. The vision tower is kept in BF16 in the checkpoint but has not been tested; the server rejects
images/videoswith HTTP 400 and runs vLLM with the vision tower disabled.- Use the bundled vLLM plugin. The quantized checkpoint does not run efficiently in plain
transformers(it can only be dequantized to BF16, which removes the memory and speed benefit). The plugin is pinned to vLLM 0.28.x; it uses vLLM pooling internals that may change in later versions.- Hardware: NVIDIA Ada, Hopper or Blackwell (SM 8.9+: L4, L40S, RTX 40/50-series, H100/H200, B200).
- Tested on one setup only: 2x RTX 5070 Ti 16 GB (SM 12.0, PCIe, no NVLink), vLLM 0.28.0, CUDA 13.
- Evaluated on 1,400 records from 7 public tasks with our own harness (below), not on Cloudflare's Decision Index. Expect a similar but not identical picture on your data; check your own workload before relying on it.
Quickstart: SystemOne API server
hf download kurcontko/clef-flash-FP8-Dynamic --local-dir clef-flash-FP8-Dynamic
docker run --gpus all --ipc=host -p 8000:8000 -v $PWD/clef-flash-FP8-Dynamic:/model:ro \
--entrypoint bash vllm/vllm-openai:v0.28.0 -c \
"cp -r /model/vllm_plugin /tmp/p && pip install -q /tmp/p && clef-systemone --model /model --dp 1"
--dp N runs one engine per GPU behind the one port (recommended over tensor parallelism on PCIe GPUs).
The server warms up its kernels before /health reports ready (about 3-4 minutes from a cold container).
curl localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"model": "clef-flash",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"}}}'
The request and response bodies are the same as the systemone() function in the original release:
answers keyed by question ID (choice with confidence and probabilities, score with the expected score
and legend, noul with the probability of true) and usage. Also served: GET /v1/models, GET /health.
Quickstart: Python (offline vLLM)
# pip install ./clef-flash-FP8-Dynamic/vllm_plugin (in an environment with vllm==0.28.*)
import sys
from transformers import AutoTokenizer
from vllm import LLM, PoolingParams
from clef_vllm import ARCHITECTURE, TASK, schema_layout
path = "./clef-flash-FP8-Dynamic"
sys.path.insert(0, path)
from joint_schema_model import encode_record
def main():
# Keep `llm` local: vLLM shuts its engine process down when the LLM object is freed.
llm = LLM(model=path, hf_overrides={"architectures": [ARCHITECTURE]}, runner="pooling",
enable_prefix_caching=False, language_model_only=True, max_model_len=16384)
tokenizer = AutoTokenizer.from_pretrained(path)
record = {
"state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
"questions": {
"status": {"type": "choice", "instructions": "What is the invoice status?",
"criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."}},
"large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
},
}
encoded = encode_record(tokenizer, record)
params = PoolingParams(task=TASK, extra_kwargs={"clef_questions": schema_layout(encoded)})
output = llm.encode([{"prompt_token_ids": list(encoded.input_ids)}], pooling_params=[params], pooling_task=TASK)[0]
logits, offset = output.outputs.data.float(), 0 # every option logit of every question, in order
for question in encoded.questions:
n = len(question.option_ids)
print(question.question_id, dict(zip(question.option_ids, logits[offset:offset + n].softmax(-1).tolist())))
offset += n
if __name__ == "__main__":
main()
Pass a list of prompts and a matching list of PoolingParams to score many records in one call; vLLM batches them.
Results
Accuracy and agreement with BF16
1,400 records, 200 per task, from held-out splits: BANKING77 (test), ARC-Challenge (test), MMLU (test),
HellaSwag (validation), ANLI R3 (test), BoolQ (validation), and a multi-question ANLI R1 set (3 questions per
record: choice + true/false + score). Accuracy is against gold labels. Agree is top-1 agreement with the
BF16 model and KL is mean KL(BF16 ‖ quantized) per question. BF16 is the original release run with
transformers 5.10.2; the quantized rows were measured in vLLM with real FP8/FP4 kernels.
| Task | BF16 acc | FP8 acc | FP8 agree | FP8 KL | NVFP4 acc | NVFP4 agree | NVFP4 KL |
|---|---|---|---|---|---|---|---|
| BANKING77 (77-way intent) | 96.5 | 97.0 | 99.5 | 0.0005 | 97.5 | 98.0 | 0.013 |
| ARC-Challenge | 99.5 | 99.5 | 100.0 | 0.0006 | 99.0 | 99.5 | 0.009 |
| MMLU | 94.5 | 95.0 | 99.5 | 0.0035 | 93.5 | 95.5 | 0.026 |
| HellaSwag | 99.0 | 99.0 | 100.0 | 0.0003 | 99.0 | 100.0 | 0.002 |
| ANLI R3 | 50.0 | 49.5 | 97.5 | 0.0028 | 49.0 | 86.5 | 0.041 |
| BoolQ | 87.5 | 88.5 | 99.0 | 0.0026 | 87.5 | 98.0 | 0.018 |
| Multi-question ANLI R1 | 73.0 | 74.0 | 98.0 | 0.0023 | 71.5 | 90.3 | 0.034 |
| All | 84.1 | 84.6 | 98.8 | 0.0019 | 83.6 | 94.3 | 0.024 |
Accuracy differences of under 1 point are within noise at 200 records per task. The BF16 model run through the same
plugin in vLLM agrees with transformers on 99.8% of questions (KL 0.0001), so the plugin itself adds no error.
NVFP4 moves most on hard, near-chance cases (ANLI): treat its confidences on borderline inputs with care.
Throughput
Same 1,400 records (717k tokens, ~512 per record), RTX 5070 Ti 16 GB GPUs.
| Setup | GPUs | Records/s | Tokens/s | vs BF16 transformers |
|---|---|---|---|---|
BF16, transformers 5.10.2 (original release code, model split over 2 GPUs) |
2 | 6.1 | 3,113 | 1.0x |
| BF16, vLLM + plugin, tensor parallel 2 | 2 | 13.0 | 6,633 | 2.1x |
| FP8, vLLM + plugin | 1 | 19.5 | 9,986 | 3.2x |
| NVFP4, vLLM + plugin | 1 | 39.8 | 20,359 | 6.5x |
FP8, one engine per GPU (--dp 2) |
2 | ~39 | ~20,000 | 6.4x |
NVFP4, one engine per GPU (--dp 2) |
2 | ~81 | ~41,000 | 13x |
Through the HTTP server with 64 concurrent clients: NVFP4 on 1 GPU 39.7 req/s (p50 0.95 s, p95 5.1 s), NVFP4 on 2 GPUs 78.7 req/s; with 8 concurrent clients on 2 GPUs, p50 latency is 132 ms. On PCIe GPUs, tensor parallelism was slower than one engine per GPU (FP8 TP2: 8,605 tok/s; NVFP4 TP2: 10,850 tok/s).
Quantization details
- Tool: llm-compressor 0.14.0 (compressed-tensors 0.19.0),
QuantizationModifier(targets="Linear", scheme="FP8_DYNAMIC"). - Scheme: FP8 (E4M3) W8A8: per-output-channel static weight scales, per-token dynamic activation scales.
- Calibration: None needed (weight scales from the weights, activation scales computed per token at runtime).
- Kept in BF16:
lm_head(the joint head reads its rows as lexical option embeddings, so it is never quantized), the vision tower (model.visual.*), the Gated DeltaNet gate projectionslinear_attn.in_proj_a/in_proj_b(32 outputs each), embeddings, norms, and the joint schema head (clef_head/). - Gated DeltaNet fused scale: vLLM fuses
linear_attn.in_proj_qkvandin_proj_zinto one GEMM, which needs a shared NVFP4 weight global scale. llm-compressor 0.14 only shares scales for q/k/v and gate/up, soscripts/quantize.pyadds("in_proj_qkv", "in_proj_z")to its fused groups. Without this, vLLM applies the larger scale to both halves and the qkv weights come out scaled by about 0.67 (we measured 92.9% instead of 94.3% agreement). Anyone re-quantizing Qwen3.5-family models to NVFP4 for vLLM should do the same. recipe.yamlis the exact recipe llm-compressor saved.
How the vLLM plugin works
vllm_plugin/ (pip install-able, package clef-vllm) registers ClefFlashForDecision, a subclass of vLLM's
Qwen3_5ForConditionalGeneration that runs as a pooling model. Its pooler collects each request's final hidden
states across chunked prefill and runs the original JointSchemaHead (from joint_schema_model.py) on the full
prompt. Each request carries its schema layout (question and option token spans from encode_record) in
PoolingParams.extra_kwargs["clef_questions"]; the output is one flat vector of every option logit. Prefix caching
must be off. Under tensor parallelism the vocab-sharded lm_head is gathered once on rank 0 for the head's lookups.
Files
| Path | Purpose |
|---|---|
model.safetensors, config.json, recipe.yaml |
Quantized backbone (12 GB, compressed-tensors format) |
clef_head/ |
Joint schema head, BF16, unchanged from the original release (in a subfolder so vLLM does not load it as backbone weights) |
joint_schema_model.py |
Original release code: encode_record, systemone_answer, the head module (unchanged) |
tokenizer*.json, chat_template.jinja, processor_config.json, generation_config.json |
Tokenizer and processor (unchanged) |
vllm_plugin/ |
vLLM plugin and clef-systemone server |
scripts/ |
build_data.py (calibration/eval sets) and quantize.py (how this checkpoint was made) |
See also kurcontko/clef-flash-NVFP4 (2x faster again, about -0.5 points).
License and attribution
Apache-2.0, following Cloudflare/clef-flash (itself post-trained from Qwen/Qwen3.5-9B). All credit for the model goes to its original authors; this repository only adds the quantization, the vLLM plugin and the server. Not an official Cloudflare release.
- Downloads last month
- 7