mlx-community/clef-flash-8bit

Cloudflare/clef-flash converted to MLX (8-bit) for Apple Silicon.

Clef turns a state (text, JSON, images, or video) plus a schema of typed questions into a probability for every allowed option, in a single forward pass. It is not a chat model — mlx_vlm.generate, mlx_lm.generate, and LM Studio will load the backbone but produce meaningless text. Use the bundled clef_mlx.py loader, which runs the backbone and the joint schema head.

Usage

pip install mlx-vlm huggingface_hub   # no torch needed
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("mlx-community/clef-flash-8bit")
sys.path.insert(0, path)
import clef_mlx

model = clef_mlx.load(path)
response = model.systemone({
    "model": "clef-flash",
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle the message?",
            "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
        },
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
        "outage": {"type": "noul", "instructions": "Is a service down?"},
    },
})
print(response["answers"])

Images (PIL) and videos (frame arrays) go in images / videos, as in the original:

from PIL import Image
model.predict({
    "state": {"task": "Review the attached receipt."},
    "images": [Image.open("receipt.jpg")],
    "questions": {"legible": {"type": "noul", "instructions": "Is the receipt total legible?"}},
})

See the original model card for the input format, question types, and benchmarks.

Conversion

  • Backbone: mlx_vlm.convert -q --q-bits 8 --q-group-size 64 (vision tower kept in bf16).
  • Joint schema head: joint_head.safetensors copied unchanged (bf16) and run by clef_mlx.py.
  • processor_config.json is the original from Cloudflare/clef-flash; prompt/token layout matches the reference joint_schema_model.py exactly (images and video).

Parity vs. official PyTorch implementation (bf16)

Inputs Top answer agrees Max abs Δprob
Text (4 records, 10 questions) 10/10 0.006
Images + video (5 records, 9 questions) 9/9 0.032

Measured on an M5 Max (128 GB). Small spot-check, not a full benchmark run.

Quality check: Decision Index (sampled)

This 8-bit model was not run on the Decision Index sample. The versions on either side of it were, on identical rows:

Decision Index (sample) Median latency
MLX bf16 55.63 339 ms
MLX 4-bit (mlx-community/clef-flash-4bit) 54.65 (96.4% same top answer as bf16) 311 ms
Cloudflare published (full suite) 57.07

8-bit tracks the PyTorch reference within 0.006 probability on spot checks, so expect it to score close to bf16.

Method: Decision Index 0.2.1 kit (suite rebuilt byte-identical), stratified 2,000-request sample across all 44 benchmarks (suite sample --n 2000), engine = clef_mlx.py with max_length=16384 and no truncation (over-length requests are refused and count as wrong; 15 of 2,000, mostly BRIGHT). The index is computed from each benchmark's native metric on the sampled rows with the kit's chance correction and weights; HLE and iSarcasmEval are set to 0 to match how Cloudflare's published run is scored. With ~40 rows per benchmark, per-benchmark numbers are noisy (±10+ pts) — only the index is meaningful.

License

Apache-2.0, following Cloudflare/clef-flash.

Downloads last month
8
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/clef-flash-8bit

Finetuned
Qwen/Qwen3.5-9B
Quantized
(18)
this model