Winnow-E4B

Choose the target model from the download links in this card or use the Winnow Quickstart. The Hub's automatic model-size/architecture summary currently describes the small Gemma-4-E4B-IT-Assistant-BF16.gguf file (172 MB), which is an optional MTP assistant and requires the matching target model. It is not the Winnow target. Both the target and assistant have BF16 files here, so use the explicit target filename or Winnow preset instead of relying on the generic :BF16 snippet. Generic Hub llama.cpp snippets do not provide Winnow's /v1/systemone API.

A compact Gemma 4 fine-tune for typed decisions that also supports ordinary chat and image input.

Winnow-E4B takes a shared state and a set of questions, then assigns probabilities to the answers each question supplies. The decision path reads candidate-answer logits without generating an explanation. It is designed for tasks such as routing a request, choosing an action, checking a condition, or rating urgency when the possible answers are known. The same merged model weights remain available for regular chat; image input uses the matching vision projector.

Winnow-E4B was created by EldanRing from Gemma 4 E4B IT. The downloadable GGUFs have the decision fine-tune merged into the language model. They require no separate adapter or base-model download.

Run the model Β· Inference code Β· API reference Β· Evaluation details

How decisions work

The Winnow inference server exposes /v1/systemone. A request provides one state and any combination of these question types:

Type What the caller supplies What the server returns
noul A yes/no question Probability of true
choice Named options with optional descriptions Selected option, option probabilities, confidence
score An ordered list of levels Level probabilities, expected score, confidence

The server shares the state prefill across questions, reuses a matching cached prefix when possible, and evaluates only the verified answer-token candidates for each question. Questions can be processed in parallel branches or waves without loading another copy of the model. Ordinary /v1/chat/completions uses the normal vocabulary head and chat template, including streaming and image messages when the projector is loaded.

For example, one request can ask both whether a customer wants a refund and which team should handle the ticket:

{
  "model": "Winnow-E4B",
  "state": {"ticket": "I was charged twice and need a refund."},
  "questions": {
    "refund": {
      "type": "noul",
      "instructions": "Does the customer request a refund?"
    },
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this ticket?",
      "criteria": {
        "billing": "Payments and refunds",
        "technical": "Software defects"
      }
    }
  },
  "winnow": {"temperature": 1.2574172017327816}
}

The quickstart shows a complete command. Candidate probabilities are conditional on the options supplied in a request. Confidence describes how concentrated the distribution is; it is not a guarantee that an answer is correct.

Model and training

This is a rank-32, alpha-64 LoRA fine-tune of the instruction-tuned Gemma 4 E4B model. The update trained language tensors; the vision and audio modules remained frozen. The adapter was merged in FP32 before conversion to Q8_0 and BF16 GGUF. Chat still has its full generation head, while the decision server projects only the requested answer rows.

The private training set combines synthetic contrastive decisions, verified labels, teacher distributions where appropriate, semantic tasks, and targeted hard cases. Training, development, calibration, and reserved tests were separated. The training data and training pipeline remain private and are not released. Public evaluation suites were used during development; the results below should be read as a comparison of deployed models, not a fresh unseen-task estimate.

Decision results

These figures evaluate the downloadable Q8 GGUF against Winnow-12B Q8 on the same questions. They report top-choice accuracy, except the local typed panel, which reports agreement with synthetic teacher labels.

Evaluation Winnow-E4B Q8 Winnow-12B Q8
JevBench public subset, 231 questions 80.52% (186/231) 85.71% (198/231)
Kev v9 clean, 1,046 questions 72.66% 81.45%
Kev v9 additional, 390 questions 55.90% 68.97%
Local typed decisions, 2,000 decisions 72.30% 70.10%

The 12B comparator uses 852/1,046 on Kev-clean, distinct from the separate 12B release campaign’s 853/1,046. These historical comparators are not interchangeable.

Winnow-12B leads on the broader public suites. E4B leads by 2.2 points on this local typed panel, which uses synthetic teacher targets rather than independently verified truth for every item. The JevBench figure is public-subset accuracy, not the official composite leaderboard score. Methods, uncertainty intervals, calibration measures, and evaluation boundaries are in the benchmark report.

Winnow-E4B and Winnow-12B Q8 decision scores on three evaluation panels

Decision quality on the same questions. The local typed panel measures agreement with synthetic teacher labels.

Tested on RTX 5070 Ti 16 GB.

In the matched 8K text workload, E4B Q8 delivered 147.8 decisions/s at 64 questions per request, versus 67.2 decisions/s for 12B Q8. These warm request medians are workload-specific. The timing table includes smaller batches and memory use.

Winnow-E4B and Winnow-12B Q8 warm decision throughput, memory, and single-decision latency

Matched text-only workload with Q8 weights and KV cache; medians of ten warm requests.

Probabilities and calibration

The server normalizes candidate logits over the options in each question. The optional winnow.temperature setting scales those logits without changing which option wins. For Q8 text decisions, 1.2574172017327816 was fitted on 778 separate calibration questions. For BF16, the separately fitted value is 1.3331553765162731; do not reuse Q8's temperature for BF16.

These temperatures were measured for the evaluated text-decision profile. Calibration can change with task, prompt length, image input, and deployment settings. Check probabilities on held-out examples from the intended application before treating them as confidence in real-world correctness.

Optional reasoning and MTP

Direct native decisions are the default. MTP drafts ordinary chat tokens; optional reasoning separately adds eligible generated context before native candidate scoring. The Winnow inference server provides verified E4B Q8 presets and the experimental text client. Direct serving needs neither an assistant nor a reasoning policy.

Build the inference server

The optional preset requires Linux/CUDA. Install the server build prerequisites and a compatible CUDA toolkit first.

git clone https://github.com/EldanRing/winnow-inference.git
cd winnow-inference
python3 scripts/build.py --backend cuda --cuda-arch 120

The E4B BF16 GGUF assistant is 171,766,688 bytes, converted without training from Google's official Gemma 4 E4B IT assistant. Use this exact E4B assistant, not the 12B draft. License and conversion attribution are in assistant documentation.

E4B Q8 8K vision+MTP

The tested e4b-q8-vision8k-mtp preset provides 8K context, vision and ordinary chat MTP4 with the matching E4B assistant. MTP and vision require additional VRAM; requirements depend on quantization, context and concurrency. Exact settings and compatibility are in setup notes. Download the Q8 model and matching projector, then launch:

python3 scripts/winnow.py download --model e4b --vision on --reasoning off --mtp on
python3 scripts/winnow.py serve --model e4b --vision on --reasoning off --mtp on
# For direct decisions and chat without MTP, use --mtp off in both commands.

E4B Q8 8K vision plus MTP4 observed peak device memory: 10,197 MiB.

The combined preset sampled 10,197 MiB peak device memory. This is an observed profile, not a sustained-capacity or vision-accuracy guarantee. No universal MTP speedup is claimed.

Experimental reasoning for one text question

e4b-calibrated50-v1 accepts one named question and a text state, preserves answer mappings, and routes on raw native T1 maxP<0.8. Completed direct/augmented probabilities use temperatures 1.2041180007310734 / 3.4209273427377678 and a 50:50 blend. These belong to this adaptive profile; they do not replace the existing direct text calibration 1.2574172017327816 or the BF16 temperature above.

Stop the vision server before starting the separate text profile:

python3 scripts/winnow.py download --model e4b --vision off --reasoning on --mtp on
python3 scripts/winnow.py serve --model e4b --vision off --reasoning on --mtp on
# For reasoning without MTP, use --mtp off for download, serve and decide.

In a second terminal from the same checkout:

python3 scripts/winnow.py decide --model e4b --reasoning on --mtp on \
  --input examples/adaptive-decision.json

The same model generates ordinary-template context; no separate thinking mode is used. Only complete eligible generation is scored and blended. After a valid direct decision, failed/incomplete reasoning returns saved policy-calibrated direct probabilities. A failed direct call remains an error. Adaptive image, structured-state and multi-question inputs are rejected. Temperature 0 generation uses a 75-second client deadline, natural EOS within the default 8K context, and a 100-word soft instruction rather than a hard output-token cap.

In the retained 900-decision/600-group LogiQA2/HelpSteer2 comparison, source-equal Brier improved 0.524807β†’0.481402 and NLL 1.591870β†’1.385153; agreement changed +0.083pp with a wide interval. Calibration and generated context are combined, so this does not isolate causal reasoning benefit or establish Boolean-transfer, unseen-data,1 pp preservation or adaptive-image coverage.

E4B historical reasoning corrects 15 and regresses 8 Jev answers; 78 and 33 Kev answers; 101 and 158 Typed agreements. Direct/adaptive rates 80.52/83.55,72.66/76.96,72.35/69.50%.

Historical E4B public-panel results: all corrections and regressions are shown. Net change is βˆ’5 across 3,277 decisions; Typed measures teacher agreement.

This figure uses a historical text-only raw T1/maxP<0.8/equal-blend experiment, different from the calibrated route and the release direct settings. Reasoning gained on Jev/Kev-clean and lost on Typed. Its E4B Typed direct panel was 1,447/2,000, separate from the release-local Typed table above. Mean mandatory direct HTTP inside the workflow versus full router time was 37β†’234ms over 3,277 decisions, including non-routed cases; HTTP excludes router overhead. These are not isolated production latency or an MTP speed benefit. Previously observed inputs and teacher labels limit generalization claims.

E4B historical workflow means: mandatory direct HTTP 37 ms; full adaptive router 234 ms across 3277 decisions. Direct HTTP is included in router timing.

Historical workflow means across 3,277 decisions, including bypasses: direct HTTP 37ms, full router 234ms. Direct HTTP is included in router time. These are not independent production arms or an isolated MTP speedup.

Presets provide defaults. Override --context (for example 16k), cache, batch and native branches where supported; choose vision, MTP and reasoning with explicit on/off flags. Run python3 scripts/winnow.py presets for short names and estimated memory. Custom settings do not inherit measured calibration or performance claims. Optional serving is tested on Linux/CUDA; Mac, multi-GPU and BF16 MTP/adaptive use are not validated. Adaptive vision calibration is not established. Direct 64K vision is a separate supported profile.

Downloads and deployment

File Role Size
Winnow-E4B-Q8_0.gguf Recommended, fully measured language model 8.01 GB / 7.46 GiB
Winnow-E4B-BF16.gguf Optional higher-precision language model 15.05 GB / 14.02 GiB
mmproj-Winnow-E4B.gguf Matching F16 projector for image input 990 MB / 0.922 GiB

Download Q8 for the measured decision, chat, and vision setup; add the projector for images. BF16 passed 8K text-decision serving but was not tested with images or populated 64K context. On reserved text tests it showed no consistent quality advantage over Q8. Verify downloads.

Q8 configuration Observed peak device memory Suggested GPU capacity
8K text-only decision workload 8.56 GiB 10 GB+ VRAM
64K vision and chat operational probes 10.72 GiB across the probes 12 GB+ VRAM

Capacity suggestions are inferred from measured usage, not qualification of every GPU in those tiers. Six cached questions over a 65,023-position image state took 129–131ms after an 11.59s first request. A separate 62,431-position multimodal prompt generated 512 tokens at 71.2 tokens/s. See the benchmark report for timing scope.

Winnow-E4B and 12B, side by side: Q8 GPU memory, chat generation, and E4B cached 64K decision timing

The 8K text workload was matched; 64K vision/chat used different images and question counts. Capacity suggestions are inferred from measured memory.

Scope and limitations

  • The decision path is intended for supplied answer options. Open-ended responses use the regular chat endpoint.
  • General chat quality, real-image vision accuracy, and broad long-context reasoning have not been established by these probes. The image test used a synthetic status panel. Vision and audio modules were not fine-tuned for this release.
  • The public suites were known during targeted data design. The local typed panel measures agreement with synthetic teacher labels. Neither result establishes accuracy for every downstream task.
  • Probability calibration was fitted on text decisions. It needs validation for a new domain or multimodal workload.
  • The 64K and vision measurements apply to Q8 with the matching projector. BF16 was evaluated only in the text profile.

Credits and license

Winnow-E4B is an independent fine-tune by EldanRing of Google DeepMind's Gemma 4 E4B IT, released under Apache 2.0. See LICENSE and NOTICE. The separate inference code builds on llama.cpp and preserves its MIT license. Jev-style describes the typed-decision interface; Winnow is not affiliated with or endorsed by TypeSafe, Google, or llama.cpp.

Downloads last month
28,414
GGUF
Model size
78M params
Architecture
gemma4-assistant
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for EldanRing/Winnow-E4B

Finetuned
(398)
this model

Spaces using EldanRing/Winnow-E4B 3

Collection including EldanRing/Winnow-E4B