Decision models experiments
System One decision models, data and benchmarks used in https://github.com/baobab-tech/decision-models-experiments
Viewer • Updated • 193k • 55Note Our evaluation set: 1,420 development evaluation reports. Document labels (silver: GLM-5.3-Flash; plus pipeline labels) and 190k tagged excerpts. Used by experiments 01 and 02.
Jev Decision Index
🔬440Benchmarks and news on various repros of TypeSafe's Jev
Note Decision Index: community leaderboard scoring open decision models and Jev on one suite.
LocalLLaMA/typed-decisions
Viewer • Updated • 3.2k • 26.4k • 108Note typed-decisions benchmark: Choice/Score/Noul items with Decider 1, d1 and Jev results.
fastino/fast-decisions
Viewer • Updated • 1.7k • 1.43k • 23Note Fastino's 17-domain decision suite (dev split public).
deepseek-ai/DeepSeek-V4.1-Flash
Image-Text-to-Text • 763B • Updated • 869k • • 4.1kNote LLM baseline and judge in experiment 01, via HF Inference Providers.
fastino/GLiNER2.5-Decide
Token Classification • 0.5B • Updated • 59.9k • 372Note DeBERTa-v3 encoder, 340M, Apache-2.0. Native multi-label; runs on CPU.
convaiinnovations/laya
Text Classification • 0.4B • Updated • 11.7k • 5.19kNote Laya: ModernBERT decision model, 421M (322M multilingual).
heman10x/rlcd-modernbert-151m
Text Classification • 0.2B • Updated • 28.9k • 38Note openJev Verdict: ModernBERT-base, 151M; trains on Apple MPS.
MaziyarPanahi/ModernJEV-Decide-Preview
Text Classification • 0.1B • Updated • 149 • 12Note ModernBERT-base, 150M, trained on 60k agent decisions.
wfzyx/von
Zero-Shot Classification • 0.4B • Updated • 88.9k • 54Note Von: ModernBERT-large decision head, 395M.
hotchpotch/bekko-system-one-v0-400m
Updated • 5Note bekko: small encoder decision model (17M–395M family).
BarraHome/Decision-Jef-0.1
Text Classification • 0.3B • Updated • 486 • 2Note Decision-Jef: 307M encoder decision model.
jaredpalmer/kev-0.8b
Text Classification • Updated • 13.2k • 39Note Kev: Qwen3.5 + LoRA + pointer head, Jev-compatible server; MLX on Mac.
jaredpalmer/kev-4b
Text Classification • Updated • 17.3k • 104Note Kev-4B: Qwen3.5-4B, Jev-compatible, MLX on Mac.
jaredpalmer/kev-9b
Text Classification • Updated • 4.56k • 67Note Kev-9B: Qwen3.5-9B.
jaredpalmer/kev-27b
Text Classification • 26B • Updated • 1.11k • 32Note Kev-27B v2: full fine-tune of Qwen3.8-27B.
alibiserikbay/JevK5
Text Generation • 4B • Updated • 9.95k • 17Note JevK5: Qwen3.5-4B with LoRA; near Jev on JevBench.
alibiserikbay/JevK5-GGUF
2B • Updated • 13.2k • 1Note JevK5 GGUF for llama.cpp (Metal).
Mapika/decider-2b
Text Classification • 2B • Updated • 305k • 101Note decider: Qwen3.5 Base decision models, 0.8B–35B.
Mapika/decider-4b
Text Classification • 4B • Updated • 23.9k • 19Note decider-4b: ranks near Jev on JevBench.
togethercomputer/Tev1-4B-experimental
Text Generation • 5B • Updated • 3.54k • 29Note Tev1: Together's Qwen3.5-4B decision model; up to ~24 options.
internlm/Intern-Decision-4B
Image-Text-to-Text • 5B • Updated • 2.3k • 76Note Intern-Decision: Qwen3.5-4B decision model.
Contrastive-LM/CLM-v0.1-8B
Text Ranking • Updated • 3.72k • 721Note CLM-8B: contrastive dual encoder on Qwen3-8B; cached option embeddings.
autotrust/JEV-27B
Text Classification • 27B • Updated • 134k • 46Note autotrust JEV-27B: Qwen3.8-27B decision model.
StrandsAgents/strands-decider-2B-hobson-v19
Text Classification • Updated • 53Note Strands Decider: 2B decision model for agent tool calls.
dwidlee/systemone-lite-0.5b
Text Generation • 0.5B • Updated • 1.89kNote systemone-lite: Qwen2.5-0.5B fine-tune.
jinghao1632/bit-jev-2b-distilled
2B • Updated • 2.43kNote bit-jev: 2B model distilled from Jev.
akhilaaa3/Jev-Omni
Text Classification • 12B • Updated • 3.73k • 366Note Jev-Omni: Gemma 4 12B, multimodal decisions.
ollaya-dev/cygnet
Text Classification • Updated • 1Note Cygnet: Gemma-4-12B; top of JevBench v1.5.4 (73.7).
perplexity-ai/pplx-decider-v1-27b
Text Classification • 26B • Updated • 988 • 81Note pplx-decider: Perplexity's 27B decision model.
Cloudflare/clef
Image-Text-to-Text • 27B • Updated • 5.42k • 1.3kNote Clef: Cloudflare's decision model.
prism-ml/Ternary-Bonsai-2-27B-gguf
Text Generation • 27B • Updated • 4.12M • 2.43kNote Ternary Bonsai 2 27B (Qwen3.8-27B), base of Bonsai-Llama-Jev.