winnow for Ollaya
Ollaya package of EldanRing/Winnow-12B and EldanRing/Winnow-E4B by EldanRing. Ollaya runs open decision models locally, the way Ollama runs LLMs: typed questions in, calibrated answers out, behind a TypeSafe-compatible API.
ollaya run winnow
What is in this repository
This repository holds only the files Ollaya derives, with no weights. The model is the authors' own GGUF
file: ollaya pull downloads it from their repository, unmodified and pinned to a commit, verifies its
sha256, and Ollaya runs it on llama.cpp.
| Tag | The authors' GGUF | Files |
|---|---|---|
winnow:12b |
EldanRing/Winnow-12B@b6ac22b gguf/Winnow-12B-Q8_0.gguf |
12b/decision.json, 12b/calibration.json |
winnow:e4b |
EldanRing/Winnow-E4B@734302f gguf/Winnow-E4B-Q8_0.gguf |
e4b/decision.json, e4b/calibration.json |
winnow:12b-vision |
EldanRing/Winnow-12B@b6ac22b gguf/Winnow-12B-Q8_0.gguf + gguf/mmproj-Winnow-12B.gguf |
12b-vision/decision.json, 12b-vision/calibration.json |
winnow:e4b-vision |
EldanRing/Winnow-E4B@734302f gguf/Winnow-E4B-Q8_0.gguf + gguf/mmproj-Winnow-E4B.gguf |
e4b-vision/decision.json, e4b-vision/calibration.json |
Each tag has decision.json (the prompt, the option labels Ollaya reads and llama.cpp's settings) and
calibration.json (temperatures). A vision tag also pulls the authors' vision projector from the same revision, which reads the images.
Parity
Ollaya's runner matches stock llama-server of the pinned build (b11146) on the same GGUF and the same device: 505 text questions per model, every decision the same, option logits within 1.3e-5 and probabilities within 3.0e-6, on CUDA (RTX 4090, Linux and Windows) and, for e4b, the x86-64 CPU. With the vision projector, 65 image questions per model, every decision the same, option logits within 1.2e-5 (RTX 4070, RTX 5090 and the x86-64 CPU).
License
Same as the upstream model (Apache-2.0). Ollaya itself is Apache-2.0.