Instructions to use EldanRing/Winnow-E2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use EldanRing/Winnow-E2B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-E2B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-E2B:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-E2B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-E2B:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf EldanRing/Winnow-E2B:BF16 # Run inference directly in the terminal: ./llama-cli -hf EldanRing/Winnow-E2B:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf EldanRing/Winnow-E2B:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf EldanRing/Winnow-E2B:BF16
Use Docker
docker model run hf.co/EldanRing/Winnow-E2B:BF16
- LM Studio
- Jan
- vLLM
How to use EldanRing/Winnow-E2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "EldanRing/Winnow-E2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EldanRing/Winnow-E2B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/EldanRing/Winnow-E2B:BF16
- Ollama
How to use EldanRing/Winnow-E2B with Ollama:
ollama run hf.co/EldanRing/Winnow-E2B:BF16
- Unsloth Desktop
- Docker Model Runner
How to use EldanRing/Winnow-E2B with Docker Model Runner:
docker model run hf.co/EldanRing/Winnow-E2B:BF16
- Lemonade
How to use EldanRing/Winnow-E2B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull EldanRing/Winnow-E2B:BF16
Run and chat with the model
lemonade run user.Winnow-E2B-BF16
List all available models
lemonade list
- Atomic Chat
Winnow-E2B
Winnow-E2B is an EldanRing decision fine-tune of Gemma 4 E2B IT. It scores supplied answer options through Winnow's /v1/systemone API and supports ordinary text chat. Image input uses the matching vision projector.
The Q8 and BF16 GGUFs contain the same merged fine-tune. No separate adapter or base-model download is needed.
Inference code · Evaluation details · Downloads
Decision results
Direct decisions score the supplied candidates. Adaptive reasoning generates context for selected questions, then blends direct and augmented probabilities. The results below use the Q8 target on an RTX 5070 Ti: F16 target K/V cache, an 8K text context, MTP4, a raw maximum-probability gate below 0.99, and a 50:50 blend.
| Panel | Direct | Adaptive reasoning |
|---|---|---|
| JevBench public, verified labels | 175/231 (75.76%) | 202/231 (87.45%) |
| Kev v9 clean, verified labels | 728/1,046 (69.60%) | 851/1,046 (81.36%) |
| Typed, synthetic teacher agreement | 1,237/2,000 (61.85%) | 1,366/2,000 (68.30%) |
All 3,277 cases remain in the totals, including two Kev context-limit attempts that fell back to direct scoring. These frozen panels were used during development. Jev and Kev use verified labels; Typed measures agreement with synthetic teacher labels.
The comparison uses Q8 targets, F16 target K/V cache, and MTP4. Reasoning prompts, policies, samplers, and batch sizes differ by model. E4B's seven-agreement Typed decline is shown to scale with a downward marker. See evaluation methods for exact settings.
Mean serial adaptive reasoning time was 1.356 s on Jev, 1.212 s on Kev, and 1.951 s on Typed. Typed timing excludes 16 cases affected by GPU overlap; all quality results remain included. See timing definitions for the measurement boundaries.
Downloads
| File | Role | Size |
|---|---|---|
| Winnow-E2B-Q8_0.gguf | Q8 target used for the reported results | 4.95 GB |
| Winnow-E2B-BF16.gguf | Higher-precision target | 9.31 GB |
| mmproj-Winnow-E2B.gguf | F16 vision projector | 986 MB |
| Gemma-4-E2B-IT-Assistant-BF16.gguf | Optional MTP draft assistant | 170 MB |
Text-only use needs one target; images also need the matching E2B projector. Download the exact filenames and verify them against SHA256SUMS.
Running the model
The quickstart covers installation and launch commands. Use /v1/systemone for decisions and /v1/chat/completions for ordinary chat and images.
F16 is the default and recommended target K/V cache. Cache precision is separate from the GGUF weight format; use --cache q8_0 for an explicit Q8 override. The pinned MTP assistant uses the target's shared K/V cache.
MTP uses the bundled BF16 conversion of Google's official E2B IT assistant. Follow the assistant download instructions; no conversion is needed. The assistant is a separate draft model, distinct from the BF16 target. Direct decisions and ordinary chat need only the target. Exact provenance is in the release manifest.
Candidate probabilities are normalized over the supplied options. Confidence describes concentration among those options, not a guarantee of correctness.
License and attribution
Gemma 4 E2B IT and the official E2B IT assistant are by Google DeepMind under Apache License 2.0. Winnow modifications are by EldanRing. See LICENSE and NOTICE. The bundled assistant includes upstream attribution. Inference code is separately licensed and is not included in this model package. No Google endorsement is implied.
- Downloads last month
- 580
8-bit
16-bit
