clef-flash-FP8-Dynamic

Unofficial FP8_DYNAMIC quantization of Cloudflare/clef-flash, a 9B decision model that scores every option of every typed question (choice / score / true-false) in one forward pass. This repo is not affiliated with or endorsed by Cloudflare.

It ships with a small vLLM plugin and a Jev/SystemOne-compatible HTTP server (POST /v1/systemone), because stock vLLM cannot run Clef's custom joint-schema head. In short: near-lossless (98.8% top-1 agreement with BF16) and 1.5x BF16's two-GPU throughput on a single GPU.

Read before using

  • Text and JSON state only. The vision tower is kept in BF16 in the checkpoint but has not been tested; the server rejects images / videos with HTTP 400 and runs vLLM with the vision tower disabled.
  • Use the bundled vLLM plugin. The quantized checkpoint does not run efficiently in plain transformers (it can only be dequantized to BF16, which removes the memory and speed benefit). The plugin is pinned to vLLM 0.28.x; it uses vLLM pooling internals that may change in later versions.
  • Hardware: NVIDIA Ada, Hopper or Blackwell (SM 8.9+: L4, L40S, RTX 40/50-series, H100/H200, B200).
  • Tested on one setup only: 2x RTX 5070 Ti 16 GB (SM 12.0, PCIe, no NVLink), vLLM 0.28.0, CUDA 13.
  • Evaluated on 1,400 records from 7 public tasks with our own harness (below), not on Cloudflare's Decision Index. Expect a similar but not identical picture on your data; check your own workload before relying on it.

Quickstart: SystemOne API server

hf download kurcontko/clef-flash-FP8-Dynamic --local-dir clef-flash-FP8-Dynamic

docker run --gpus all --ipc=host -p 8000:8000 -v $PWD/clef-flash-FP8-Dynamic:/model:ro \
  --entrypoint bash vllm/vllm-openai:v0.28.0 -c \
  "cp -r /model/vllm_plugin /tmp/p && pip install -q /tmp/p && clef-systemone --model /model --dp 1"

--dp N runs one engine per GPU behind the one port (recommended over tensor parallelism on PCIe GPUs). The server warms up its kernels before /health reports ready (about 3-4 minutes from a cold container).

curl localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
  "model": "clef-flash",
  "state": "Our checkout started returning errors and orders are blocked.",
  "questions": {
    "department": {"type": "choice", "instructions": "Which team should handle the message?",
                   "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}},
    "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
    "outage": {"type": "noul", "instructions": "Is a service down?"}}}'

The request and response bodies are the same as the systemone() function in the original release: answers keyed by question ID (choice with confidence and probabilities, score with the expected score and legend, noul with the probability of true) and usage. Also served: GET /v1/models, GET /health.

Quickstart: Python (offline vLLM)

# pip install ./clef-flash-FP8-Dynamic/vllm_plugin   (in an environment with vllm==0.28.*)
import sys
from transformers import AutoTokenizer
from vllm import LLM, PoolingParams
from clef_vllm import ARCHITECTURE, TASK, schema_layout

path = "./clef-flash-FP8-Dynamic"
sys.path.insert(0, path)
from joint_schema_model import encode_record


def main():
    # Keep `llm` local: vLLM shuts its engine process down when the LLM object is freed.
    llm = LLM(model=path, hf_overrides={"architectures": [ARCHITECTURE]}, runner="pooling",
              enable_prefix_caching=False, language_model_only=True, max_model_len=16384)
    tokenizer = AutoTokenizer.from_pretrained(path)
    record = {
        "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
        "questions": {
            "status": {"type": "choice", "instructions": "What is the invoice status?",
                       "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."}},
            "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
        },
    }
    encoded = encode_record(tokenizer, record)
    params = PoolingParams(task=TASK, extra_kwargs={"clef_questions": schema_layout(encoded)})
    output = llm.encode([{"prompt_token_ids": list(encoded.input_ids)}], pooling_params=[params], pooling_task=TASK)[0]

    logits, offset = output.outputs.data.float(), 0   # every option logit of every question, in order
    for question in encoded.questions:
        n = len(question.option_ids)
        print(question.question_id, dict(zip(question.option_ids, logits[offset:offset + n].softmax(-1).tolist())))
        offset += n


if __name__ == "__main__":
    main()

Pass a list of prompts and a matching list of PoolingParams to score many records in one call; vLLM batches them.

Results

Accuracy and agreement with BF16

1,400 records, 200 per task, from held-out splits: BANKING77 (test), ARC-Challenge (test), MMLU (test), HellaSwag (validation), ANLI R3 (test), BoolQ (validation), and a multi-question ANLI R1 set (3 questions per record: choice + true/false + score). Accuracy is against gold labels. Agree is top-1 agreement with the BF16 model and KL is mean KL(BF16 ‖ quantized) per question. BF16 is the original release run with transformers 5.10.2; the quantized rows were measured in vLLM with real FP8/FP4 kernels.

Task BF16 acc FP8 acc FP8 agree FP8 KL NVFP4 acc NVFP4 agree NVFP4 KL
BANKING77 (77-way intent) 96.5 97.0 99.5 0.0005 97.5 98.0 0.013
ARC-Challenge 99.5 99.5 100.0 0.0006 99.0 99.5 0.009
MMLU 94.5 95.0 99.5 0.0035 93.5 95.5 0.026
HellaSwag 99.0 99.0 100.0 0.0003 99.0 100.0 0.002
ANLI R3 50.0 49.5 97.5 0.0028 49.0 86.5 0.041
BoolQ 87.5 88.5 99.0 0.0026 87.5 98.0 0.018
Multi-question ANLI R1 73.0 74.0 98.0 0.0023 71.5 90.3 0.034
All 84.1 84.6 98.8 0.0019 83.6 94.3 0.024

Accuracy differences of under 1 point are within noise at 200 records per task. The BF16 model run through the same plugin in vLLM agrees with transformers on 99.8% of questions (KL 0.0001), so the plugin itself adds no error. NVFP4 moves most on hard, near-chance cases (ANLI): treat its confidences on borderline inputs with care.

Throughput

Same 1,400 records (717k tokens, ~512 per record), RTX 5070 Ti 16 GB GPUs.

Setup GPUs Records/s Tokens/s vs BF16 transformers
BF16, transformers 5.10.2 (original release code, model split over 2 GPUs) 2 6.1 3,113 1.0x
BF16, vLLM + plugin, tensor parallel 2 2 13.0 6,633 2.1x
FP8, vLLM + plugin 1 19.5 9,986 3.2x
NVFP4, vLLM + plugin 1 39.8 20,359 6.5x
FP8, one engine per GPU (--dp 2) 2 ~39 ~20,000 6.4x
NVFP4, one engine per GPU (--dp 2) 2 ~81 ~41,000 13x

Through the HTTP server with 64 concurrent clients: NVFP4 on 1 GPU 39.7 req/s (p50 0.95 s, p95 5.1 s), NVFP4 on 2 GPUs 78.7 req/s; with 8 concurrent clients on 2 GPUs, p50 latency is 132 ms. On PCIe GPUs, tensor parallelism was slower than one engine per GPU (FP8 TP2: 8,605 tok/s; NVFP4 TP2: 10,850 tok/s).

Quantization details

  • Tool: llm-compressor 0.14.0 (compressed-tensors 0.19.0), QuantizationModifier(targets="Linear", scheme="FP8_DYNAMIC").
  • Scheme: FP8 (E4M3) W8A8: per-output-channel static weight scales, per-token dynamic activation scales.
  • Calibration: None needed (weight scales from the weights, activation scales computed per token at runtime).
  • Kept in BF16: lm_head (the joint head reads its rows as lexical option embeddings, so it is never quantized), the vision tower (model.visual.*), the Gated DeltaNet gate projections linear_attn.in_proj_a / in_proj_b (32 outputs each), embeddings, norms, and the joint schema head (clef_head/).
  • Gated DeltaNet fused scale: vLLM fuses linear_attn.in_proj_qkv and in_proj_z into one GEMM, which needs a shared NVFP4 weight global scale. llm-compressor 0.14 only shares scales for q/k/v and gate/up, so scripts/quantize.py adds ("in_proj_qkv", "in_proj_z") to its fused groups. Without this, vLLM applies the larger scale to both halves and the qkv weights come out scaled by about 0.67 (we measured 92.9% instead of 94.3% agreement). Anyone re-quantizing Qwen3.5-family models to NVFP4 for vLLM should do the same.
  • recipe.yaml is the exact recipe llm-compressor saved.

How the vLLM plugin works

vllm_plugin/ (pip install-able, package clef-vllm) registers ClefFlashForDecision, a subclass of vLLM's Qwen3_5ForConditionalGeneration that runs as a pooling model. Its pooler collects each request's final hidden states across chunked prefill and runs the original JointSchemaHead (from joint_schema_model.py) on the full prompt. Each request carries its schema layout (question and option token spans from encode_record) in PoolingParams.extra_kwargs["clef_questions"]; the output is one flat vector of every option logit. Prefix caching must be off. Under tensor parallelism the vocab-sharded lm_head is gathered once on rank 0 for the head's lookups.

Files

Path Purpose
model.safetensors, config.json, recipe.yaml Quantized backbone (12 GB, compressed-tensors format)
clef_head/ Joint schema head, BF16, unchanged from the original release (in a subfolder so vLLM does not load it as backbone weights)
joint_schema_model.py Original release code: encode_record, systemone_answer, the head module (unchanged)
tokenizer*.json, chat_template.jinja, processor_config.json, generation_config.json Tokenizer and processor (unchanged)
vllm_plugin/ vLLM plugin and clef-systemone server
scripts/ build_data.py (calibration/eval sets) and quantize.py (how this checkpoint was made)

See also kurcontko/clef-flash-NVFP4 (2x faster again, about -0.5 points).

License and attribution

Apache-2.0, following Cloudflare/clef-flash (itself post-trained from Qwen/Qwen3.5-9B). All credit for the model goes to its original authors; this repository only adds the quantization, the vLLM plugin and the server. Not an official Cloudflare release.

Downloads last month
7
Safetensors
Model size
9B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kurcontko/clef-flash-FP8-Dynamic

Finetuned
Qwen/Qwen3.5-9B
Quantized
(26)
this model