Instructions to use EldanRing/Winnow-E4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use EldanRing/Winnow-E4B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-E4B:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-E4B:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: ./llama-cli -hf EldanRing/Winnow-E4B:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf EldanRing/Winnow-E4B:BF16
Use Docker
docker model run hf.co/EldanRing/Winnow-E4B:BF16
- LM Studio
- Jan
- vLLM
How to use EldanRing/Winnow-E4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "EldanRing/Winnow-E4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EldanRing/Winnow-E4B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/EldanRing/Winnow-E4B:BF16
- Ollama
How to use EldanRing/Winnow-E4B with Ollama:
ollama run hf.co/EldanRing/Winnow-E4B:BF16
- Unsloth Desktop
- Docker Model Runner
How to use EldanRing/Winnow-E4B with Docker Model Runner:
docker model run hf.co/EldanRing/Winnow-E4B:BF16
- Lemonade
How to use EldanRing/Winnow-E4B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull EldanRing/Winnow-E4B:BF16
Run and chat with the model
lemonade run user.Winnow-E4B-BF16
List all available models
lemonade list
- Atomic Chat
Winnow-E4B
Choose the target model from the download links in this card or use the
Winnow Quickstart.
The Hub's automatic model-size/architecture summary currently describes the small
Gemma-4-E4B-IT-Assistant-BF16.gguf file (172 MB), which is an optional MTP
assistant and requires the matching target model. It is not the Winnow target.
Both the target and assistant have BF16 files here, so use the explicit target
filename or Winnow preset instead of relying on the generic :BF16 snippet.
Generic Hub llama.cpp snippets do not provide Winnow's /v1/systemone API.
A compact Gemma 4 fine-tune for typed decisions that also supports ordinary chat and image input.
Winnow-E4B takes a shared state and a set of questions, then assigns probabilities to the answers each question supplies. The decision path reads candidate-answer logits without generating an explanation. It is designed for tasks such as routing a request, choosing an action, checking a condition, or rating urgency when the possible answers are known. The same merged model weights remain available for regular chat; image input uses the matching vision projector.
Winnow-E4B was created by EldanRing from Gemma 4 E4B IT. The downloadable GGUFs have the decision fine-tune merged into the language model. They require no separate adapter or base-model download.
Run the model Β· Inference code Β· API reference Β· Evaluation details
How decisions work
The Winnow inference server exposes /v1/systemone. A request provides one state and any combination of these question types:
| Type | What the caller supplies | What the server returns |
|---|---|---|
noul |
A yes/no question | Probability of true |
choice |
Named options with optional descriptions | Selected option, option probabilities, confidence |
score |
An ordered list of levels | Level probabilities, expected score, confidence |
The server shares the state prefill across questions, reuses a matching cached prefix when possible, and evaluates only the verified answer-token candidates for each question. Questions can be processed in parallel branches or waves without loading another copy of the model. Ordinary /v1/chat/completions uses the normal vocabulary head and chat template, including streaming and image messages when the projector is loaded.
For example, one request can ask both whether a customer wants a refund and which team should handle the ticket:
{
"model": "Winnow-E4B",
"state": {"ticket": "I was charged twice and need a refund."},
"questions": {
"refund": {
"type": "noul",
"instructions": "Does the customer request a refund?"
},
"department": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Payments and refunds",
"technical": "Software defects"
}
}
},
"winnow": {"temperature": 1.2574172017327816}
}
The quickstart shows a complete command. Candidate probabilities are conditional on the options supplied in a request. Confidence describes how concentrated the distribution is; it is not a guarantee that an answer is correct.
Model and training
This is a rank-32, alpha-64 LoRA fine-tune of the instruction-tuned Gemma 4 E4B model. The update trained language tensors; the vision and audio modules remained frozen. The adapter was merged in FP32 before conversion to Q8_0 and BF16 GGUF. Chat still has its full generation head, while the decision server projects only the requested answer rows.
The private training set combines synthetic contrastive decisions, verified labels, teacher distributions where appropriate, semantic tasks, and targeted hard cases. Training, development, calibration, and reserved tests were separated. The training data and training pipeline remain private and are not released. Public evaluation suites were used during development; the results below should be read as a comparison of deployed models, not a fresh unseen-task estimate.
Decision results
These figures evaluate the downloadable Q8 GGUF against Winnow-12B Q8 on the same questions. They report top-choice accuracy, except the local typed panel, which reports agreement with synthetic teacher labels.
| Evaluation | Winnow-E4B Q8 | Winnow-12B Q8 |
|---|---|---|
| JevBench public subset, 231 questions | 80.52% (186/231) | 85.71% (198/231) |
| Kev v9 clean, 1,046 questions | 72.66% | 81.45% |
| Kev v9 additional, 390 questions | 55.90% | 68.97% |
| Local typed decisions, 2,000 decisions | 72.30% | 70.10% |
The 12B comparator uses 852/1,046 on Kev-clean, distinct from the separate 12B release campaignβs 853/1,046. These historical comparators are not interchangeable.
Winnow-12B leads on the broader public suites. E4B leads by 2.2 points on this local typed panel, which uses synthetic teacher targets rather than independently verified truth for every item. The JevBench figure is public-subset accuracy, not the official composite leaderboard score. Methods, uncertainty intervals, calibration measures, and evaluation boundaries are in the benchmark report.
Decision quality on the same questions. The local typed panel measures agreement with synthetic teacher labels.
Tested on RTX 5070 Ti 16 GB.
In the matched 8K text workload, E4B Q8 delivered 147.8 decisions/s at 64 questions per request, versus 67.2 decisions/s for 12B Q8. These warm request medians are workload-specific. The timing table includes smaller batches and memory use.
Matched text-only workload with Q8 weights and KV cache; medians of ten warm requests.
Probabilities and calibration
The server normalizes candidate logits over the options in each question. The optional winnow.temperature setting scales those logits without changing which option wins. For Q8 text decisions, 1.2574172017327816 was fitted on 778 separate calibration questions. For BF16, the separately fitted value is 1.3331553765162731; do not reuse Q8's temperature for BF16.
These temperatures were measured for the evaluated text-decision profile. Calibration can change with task, prompt length, image input, and deployment settings. Check probabilities on held-out examples from the intended application before treating them as confidence in real-world correctness.
Optional reasoning and MTP
Direct native decisions are the default. MTP drafts ordinary chat tokens; optional reasoning separately adds eligible generated context before native candidate scoring. The Winnow inference server provides verified E4B Q8 presets and the experimental text client. Direct serving needs neither an assistant nor a reasoning policy.
Build the inference server
The optional preset requires Linux/CUDA. Install the server build prerequisites and a compatible CUDA toolkit first.
git clone https://github.com/EldanRing/winnow-inference.git
cd winnow-inference
python3 scripts/build.py --backend cuda --cuda-arch 120
The E4B BF16 GGUF assistant is 171,766,688 bytes, converted without training from Google's official Gemma 4 E4B IT assistant. Use this exact E4B assistant, not the 12B draft. License and conversion attribution are in assistant documentation.
E4B Q8 8K vision+MTP
The tested e4b-q8-vision8k-mtp preset provides 8K context, vision and ordinary
chat MTP4 with the matching E4B assistant. MTP and vision require additional VRAM;
requirements depend on quantization, context and concurrency. Exact settings and
compatibility are in setup notes. Download the Q8 model
and matching projector, then launch:
python3 scripts/winnow.py download --model e4b --vision on --reasoning off --mtp on
python3 scripts/winnow.py serve --model e4b --vision on --reasoning off --mtp on
# For direct decisions and chat without MTP, use --mtp off in both commands.
The combined preset sampled 10,197 MiB peak device memory. This is an observed profile, not a sustained-capacity or vision-accuracy guarantee. No universal MTP speedup is claimed.
Experimental reasoning for one text question
e4b-calibrated50-v1 accepts one named question and a text state, preserves
answer mappings, and routes on raw native T1 maxP<0.8. Completed direct/augmented
probabilities use temperatures 1.2041180007310734 / 3.4209273427377678 and a 50:50
blend. These belong to this adaptive profile; they do not replace the existing
direct text calibration 1.2574172017327816 or the BF16 temperature above.
Stop the vision server before starting the separate text profile:
python3 scripts/winnow.py download --model e4b --vision off --reasoning on --mtp on
python3 scripts/winnow.py serve --model e4b --vision off --reasoning on --mtp on
# For reasoning without MTP, use --mtp off for download, serve and decide.
In a second terminal from the same checkout:
python3 scripts/winnow.py decide --model e4b --reasoning on --mtp on \
--input examples/adaptive-decision.json
The same model generates ordinary-template context; no separate thinking mode is used. Only complete eligible generation is scored and blended. After a valid direct decision, failed/incomplete reasoning returns saved policy-calibrated direct probabilities. A failed direct call remains an error. Adaptive image, structured-state and multi-question inputs are rejected. Temperature 0 generation uses a 75-second client deadline, natural EOS within the default 8K context, and a 100-word soft instruction rather than a hard output-token cap.
In the retained 900-decision/600-group LogiQA2/HelpSteer2 comparison, source-equal Brier improved 0.524807β0.481402 and NLL 1.591870β1.385153; agreement changed +0.083pp with a wide interval. Calibration and generated context are combined, so this does not isolate causal reasoning benefit or establish Boolean-transfer, unseen-data,1 pp preservation or adaptive-image coverage.
Historical E4B public-panel results: all corrections and regressions are shown. Net change is β5 across 3,277 decisions; Typed measures teacher agreement.
This figure uses a historical text-only raw T1/maxP<0.8/equal-blend experiment, different from the calibrated route and the release direct settings. Reasoning gained on Jev/Kev-clean and lost on Typed. Its E4B Typed direct panel was 1,447/2,000, separate from the release-local Typed table above. Mean mandatory direct HTTP inside the workflow versus full router time was 37β234ms over 3,277 decisions, including non-routed cases; HTTP excludes router overhead. These are not isolated production latency or an MTP speed benefit. Previously observed inputs and teacher labels limit generalization claims.
Historical workflow means across 3,277 decisions, including bypasses: direct HTTP 37ms, full router 234ms. Direct HTTP is included in router time. These are not independent production arms or an isolated MTP speedup.
Presets provide defaults. Override --context (for example 16k), cache, batch
and native branches where supported; choose vision, MTP and reasoning with
explicit on/off flags. Run python3 scripts/winnow.py presets for short names
and estimated memory. Custom settings do not inherit measured calibration or
performance claims. Optional serving is tested on Linux/CUDA;
Mac, multi-GPU and BF16 MTP/adaptive use are not validated. Adaptive vision
calibration is not established. Direct 64K vision is a separate supported profile.
Downloads and deployment
| File | Role | Size |
|---|---|---|
| Winnow-E4B-Q8_0.gguf | Recommended, fully measured language model | 8.01 GB / 7.46 GiB |
| Winnow-E4B-BF16.gguf | Optional higher-precision language model | 15.05 GB / 14.02 GiB |
| mmproj-Winnow-E4B.gguf | Matching F16 projector for image input | 990 MB / 0.922 GiB |
Download Q8 for the measured decision, chat, and vision setup; add the projector for images. BF16 passed 8K text-decision serving but was not tested with images or populated 64K context. On reserved text tests it showed no consistent quality advantage over Q8. Verify downloads.
| Q8 configuration | Observed peak device memory | Suggested GPU capacity |
|---|---|---|
| 8K text-only decision workload | 8.56 GiB | 10 GB+ VRAM |
| 64K vision and chat operational probes | 10.72 GiB across the probes | 12 GB+ VRAM |
Capacity suggestions are inferred from measured usage, not qualification of every GPU in those tiers. Six cached questions over a 65,023-position image state took 129β131ms after an 11.59s first request. A separate 62,431-position multimodal prompt generated 512 tokens at 71.2 tokens/s. See the benchmark report for timing scope.
The 8K text workload was matched; 64K vision/chat used different images and question counts. Capacity suggestions are inferred from measured memory.
Scope and limitations
- The decision path is intended for supplied answer options. Open-ended responses use the regular chat endpoint.
- General chat quality, real-image vision accuracy, and broad long-context reasoning have not been established by these probes. The image test used a synthetic status panel. Vision and audio modules were not fine-tuned for this release.
- The public suites were known during targeted data design. The local typed panel measures agreement with synthetic teacher labels. Neither result establishes accuracy for every downstream task.
- Probability calibration was fitted on text decisions. It needs validation for a new domain or multimodal workload.
- The 64K and vision measurements apply to Q8 with the matching projector. BF16 was evaluated only in the text profile.
Credits and license
Winnow-E4B is an independent fine-tune by EldanRing of Google DeepMind's Gemma 4 E4B IT, released under Apache 2.0. See LICENSE and NOTICE. The separate inference code builds on llama.cpp and preserves its MIT license. Jev-style describes the typed-decision interface; Winnow is not affiliated with or endorsed by TypeSafe, Google, or llama.cpp.
- Downloads last month
- 28,414
8-bit
16-bit





