Instructions to use mlx-community/clef-flash-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/clef-flash-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("mlx-community/clef-flash-4bit") config = load_config("mlx-community/clef-flash-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use mlx-community/clef-flash-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/clef-flash-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "mlx-community/clef-flash-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use mlx-community/clef-flash-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/clef-flash-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default mlx-community/clef-flash-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use mlx-community/clef-flash-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/clef-flash-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "mlx-community/clef-flash-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
mlx-community/clef-flash-4bit
Cloudflare/clef-flash converted to MLX (4-bit) for Apple Silicon.
Clef turns a state (text, JSON, images, or video) plus a schema of typed questions into a
probability for every allowed option, in a single forward pass. It is not a chat model —
mlx_vlm.generate, mlx_lm.generate, and LM Studio will load the backbone but produce
meaningless text. Use the bundled clef_mlx.py loader, which runs the backbone and the
joint schema head.
Usage
pip install mlx-vlm huggingface_hub # no torch needed
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("mlx-community/clef-flash-4bit")
sys.path.insert(0, path)
import clef_mlx
model = clef_mlx.load(path)
response = model.systemone({
"model": "clef-flash",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
})
print(response["answers"])
Images (PIL) and videos (frame arrays) go in images / videos, as in the original:
from PIL import Image
model.predict({
"state": {"task": "Review the attached receipt."},
"images": [Image.open("receipt.jpg")],
"questions": {"legible": {"type": "noul", "instructions": "Is the receipt total legible?"}},
})
See the original model card for the input format, question types, and benchmarks.
Conversion
- Backbone:
mlx_vlm.convert -q --q-bits 4 --q-group-size 64(vision tower kept in bf16). - Joint schema head:
joint_head.safetensorscopied unchanged (bf16) and run byclef_mlx.py. processor_config.jsonis the original from Cloudflare/clef-flash; prompt/token layout matches the referencejoint_schema_model.pyexactly (images and video).
Parity vs. official PyTorch implementation (bf16)
| Inputs | Top answer agrees | Max abs Δprob |
|---|---|---|
| Text (4 records, 10 questions) | 10/10 | 0.029 |
| Images + video (5 records, 9 questions) | 9/9 | 0.119 |
Measured on an M5 Max (128 GB). Small spot-check, not a full benchmark run.
Quality check: Decision Index (sampled)
| Decision Index (sample) | Same top answer as bf16 | Mean max abs Δp | Median latency | |
|---|---|---|---|---|
| This model (4-bit) | 54.65 | 96.4% (15,915 answers) | 0.040 | 311 ms |
| MLX bf16, same rows | 55.63 | — | — | 339 ms |
| Cloudflare published (full suite) | 57.07 |
4-bit costs about 1 index point vs bf16 on identical rows. Most of the remaining gap to the published score is already present in bf16 (sample noise and harness differences), not quantization. Agreement is lowest on low-chance many-option tasks (POP909, GPQA, CLINC150).
Method: Decision Index 0.2.1 kit (suite rebuilt byte-identical), stratified
2,000-request sample across all 44 benchmarks (suite sample --n 2000), engine = clef_mlx.py with
max_length=16384 and no truncation (over-length requests are refused and count as wrong; 15 of 2,000,
mostly BRIGHT). The index is computed from each benchmark's native metric on the sampled rows with the kit's
chance correction and weights; HLE and iSarcasmEval are set to 0 to match how Cloudflare's published run is
scored. With ~40 rows per benchmark, per-benchmark numbers are noisy (±10+ pts) — only the index is meaningful.
License
Apache-2.0, following Cloudflare/clef-flash.
- Downloads last month
- -
4-bit