Instructions to use autotrust/GEV-26B-Decide with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use autotrust/GEV-26B-Decide with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="autotrust/GEV-26B-Decide")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("autotrust/GEV-26B-Decide") model = AutoModelForMultimodalLM.from_pretrained("autotrust/GEV-26B-Decide", device_map="auto") - Notebooks
- Google Colab
- Kaggle
autotrust/GEV-26B-Decide
A decision model with adaptive thinking, on Gemma-4-26B-A4B-it (26 B parameters, ≈ 4 B active per token) — Decision Index 0.2.1: 62.48
Decision Index
| Decision Index 0.2.1 (balanced skill) | balanced raw | breadth skill | |
|---|---|---|---|
| autotrust/GEV-26B-Decide, adaptive thinking | 62.48 | 70.66 | 62.00 |
| TypeSafe Jev 1.13 (board) | 57.91 | — | — |
| area (skill) | Knowledge & Reasoning | Language | Retrieval & Classification | Tools & Automation | Arts & Taste |
|---|---|---|---|---|---|
| GEV-26B-Decide, adaptive thinking | 0.602 | 0.636 | 0.679 | 0.697 | 0.415 |
How the score was computed (our scoring with the kit's score --edition 0.2.1, not a board entry):
- Knowledge & Reasoning: all ten benchmarks with adaptive thinking (per-benchmark table under Adaptive thinking).
- Other four areas: System 1 only (thinking off), from the complete System 1 run of these weights (all 150,759
requests, 0 errors), results in
autotrust/jev-decision-index-results(runs/jev-gemma4-26b-a4b, the weights' previous name). Thinking was tried on five of their benchmarks and is not used there: it adds little to classification, retrieval and tool selection (see Outside Knowledge & Reasoning). - The board requires a median latency of at most 1,000 ms per request, measured on an RTX PRO 6000. System 1 answers in about 45 ms on a B200; adaptive thinking is far slower on the Knowledge & Reasoning benchmarks (see Latency).
Details: reports/decision_index_adaptive.json, reports/adaptive_latency_summary.json.
Overview
GEV-26B-Decide answers typed questions with a calibrated probability for every option; with thinking switched on, it thinks only when it needs to. System 1 decides in one forward pass (about 45 ms). When its leading option is uncertain, System 2 (the same backbone in Gemma-4 thinking mode) reasons over the question, and the reasoning is folded into the final probabilities. One set of weights, one vLLM engine, for text and images.
| what it does | output | |
|---|---|---|
| System 1 | typed decisions: yes/no · pick one of 2–256 options · rate 0–5, over text and images; prompts up to 256K tokens | a calibrated probability for every option, in one forward pass |
Adaptive thinking (opt-in: thinking: "auto") |
System 1 first; below 0.8 confidence, System 2 thinks and its answer is folded in | calibrated probabilities |
| System 2 | the unmodified google/gemma-4-26B-A4B-it, optionally thinking step by step, text and images |
text / reasoning |
GEV-26B-Decide was previously published as autotrust/JEV-Gemma4-26B-A4B; the weights are the same.
Two models, two organisations. TypeSafe Jev 1.13 is the hosted, closed model made by TypeSafe AI. autotrust/GEV-26B-Decide is an independent open-weights model built by AutoTrust AI; it is not affiliated with, endorsed by, or a product of TypeSafe AI.
Watch it think
Each puzzle has one correct answer. System 1 answers in one pass; with thinking: "auto", System 2 thinks when System 1
is below 0.8 confidence, and its answer is folded into the probabilities. The videos show puzzles that System 1 got wrong;
the tables below count all puzzles.
Minesweeper. Which hidden cell is certainly safe? System 1 is at chance (21.0 % against 25 %); with thinking, 86.0 %. These thoughts usually reach the 8,192-token budget, and the answer read at that point is still right most of the time.
Connect Four. Which column wins now, or stops the opponent from winning next move? 54.0 % → 99.5 %.
Wordle. Which word still fits all the colour feedback? 52.0 % → 100 %.
Sudoku. Which digit belongs in the highlighted cell? 76.0 % → 99.5 %.
One-move puzzles, 200 generated puzzles per game (threshold 0.8, budget 8,192 thinking tokens; text input, chess with the board image as well):
| game | question (options) | chance | System 1 | adaptive | puzzles that thought |
|---|---|---|---|---|---|
| Minesweeper | which hidden cell is certainly safe (1 safe cell, 3 mines) | 25.0 | 21.0 | 86.0 | 100 % |
| Wordle | which word fits all the feedback (8 words) | 12.5 | 52.0 | 100.0 | 98 % |
| Connect Four | which column wins now or blocks (legal columns) | 14.5 | 54.0 | 99.5 | 95 % |
| Sudoku | which digit goes in the cell (1–9) | 11.1 | 76.0 | 99.5 | 92 % |
| 24 game | which expression equals 24 (6 expressions) | 16.7 | 89.0 | 100.0 | 74 % |
| Maze | first step towards the exit (2–4 directions) | 48.4 | 48.5 | 68.0 | 82 % |
| Chess (Lichess puzzles) | which move mates in one (16 moves, board image + FEN) | 6.3 | 46.5 | 78.0 | — |
Whole games and perception, where thinking helps little or not at all:
| check | System 1 | adaptive |
|---|---|---|
| Snake, one game (image + positions) | 3 food in 18 steps | 11 food in 90 steps (about 21 s per step) |
| Connect Four, 6 full games against a heuristic opponent | 1 win, 5 losses | 1 win, 4 losses, 1 draw |
| 2048, one game | 1,476 points | 1,016 points |
| Flappy Bird, one game | 0 pipes (23 frames) | 0 pipes (38 frames, about 35 s per frame) |
| Quick, Draw!, 320 real sketches, 16 answers (30 / 60 / 100 % of the strokes) | 46.9 / 73.1 / 94.4 % | 40.3 / 72.8 / 95.0 % |
Thinking pays off when the answer can be checked step by step against explicit rules: logic puzzles, tactics,
constraints, arithmetic. It does not help perception (sketches), reflexes (Flappy Bird) or long-horizon play (2048, full
Connect Four games), and every thought costs seconds. Details: reports/thinking_games.json.
Adaptive thinking
- Fast distribution. System 1 returns p1 in one pass.
- Think only when uncertain. If the leading option of p1 is below the threshold (default 0.8), System 2 reasons in
Gemma-4's thinking mode over the same state, question and options. The reasoning length is yours to set, as with the
base model:
think_budgetcaps the thinking tokens (default: no cap beyond the context window; the evaluations below used 8,192). - Fold the reasoning in. When the thinking channel closes, the answer-letter distribution p2 is read in one step, and the result is p = ½ p1 + ½ p2. On its own, p2 is close to one-hot and over-confident; the equal mix keeps the reasoning's accuracy and System 1's calibration.
The threshold and the mix were chosen on 1,754 questions from six public sets that are not part of the Decision Index (test or validation splits, 300 random questions each; AQuA-RAT has 254):
| set | System 1 | adaptive | thinking on | always think |
|---|---|---|---|---|
| AQuA-RAT (math word problems) | 68.1 | 89.0 | 53.9 % | 90.2 |
| LogiQA (logical reasoning) | 55.3 | 80.3 | 63.7 % | 82.0 |
| StrategyQA (multi-hop yes/no) | 68.0 | 78.7 | 71.0 % | 79.0 |
| MedMCQA (medical) | 66.7 | 71.7 | 52.3 % | 73.7 |
| OpenBookQA (science) | 94.3 | 96.3 | 14.7 % | 95.7 |
| CommonsenseQA | 86.7 | 85.0 | 32.0 % | 83.3 |
| all 1,754 | 73.3 | 83.4 | 47.8 % | 83.8 |
Accuracy in %. Calibration is unchanged: ECE 0.035 for System 1 and 0.035 for adaptive. The adaptive mode reaches 96 % of the always-think gain while thinking on 48 % of the questions. Thinking length: median 2,282 tokens, 90th percentile 7,573; 91.6 % of the thoughts finish within the 8,192-token budget.
Decision Index, Knowledge & Reasoning
All ten benchmarks of the area (33,856 scored requests) ran through the server (POST /v1/decide, thinking: "auto",
threshold 0.8, budget 8,192 thinking tokens) and were scored with the Decision Index kit:
| benchmark | chance | System 1 | adaptive | DI skill, System 1 → adaptive | questions that thought |
|---|---|---|---|---|---|
| GPQA Diamond | 25.0 | 42.9 | 78.6 | 0.238 → 0.714 | 87 % |
| CRUXEval | 37.0 | 67.5 | 90.7 | 0.485 → 0.853 | 42 % |
| CLadder | 50.0 | 71.0 | 86.6 | 0.420 → 0.732 | 51 % |
| GSM8K | 25.0 | 97.6 | 99.1 | 0.969 → 0.989 | 5 % |
| SATA-Bench (case exact) | 1.3 | 34.2 | 35.5 | 0.334 → 0.346 | 21 % |
| MuSR | 37.1 | 67.3 | 67.6 | 0.480 → 0.484 | 51 % |
| ChessBench | 8.2 | 23.7 | 23.6 | 0.169 → 0.168 | 88 % |
| HLE | 16.4 | 8.4 | 17.8 | 0.000 → 0.016 | 88 % |
| MMLU-Pro | 11.1 | 65.0 | 84.6 | 0.607 → 0.827 | 78 % |
| BBH | 31.0 | 75.0 | 92.0 | 0.638 → 0.884 | 60 % |
Accuracy in %; "chance" is the kit's random baseline, and the Decision Index skill rescales accuracy so that chance is 0 (below chance counts as 0). The Knowledge & Reasoning area skill rises from 0.429 to 0.602.
- Real gains on science, code, causal and multi-step reasoning: GPQA Diamond, CRUXEval, CLadder, MMLU-Pro and BBH.
- HLE only returns to chance. System 1 is below chance (8.4 % against 16.4 %), and the questions where it is confident enough not to think are almost all wrong (1.6 %). Thinking brings HLE to 17.8 %, the level of guessing, so its skill stays near 0.
- Chess does not benefit. Two thirds of the ChessBench thoughts hit the 8,192-token budget; the answer read from a truncated thought is no better than System 1 (17.0 % against 17.7 % on those questions), while finished thoughts gain a little (19.8 % → 22.7 %). Answers read from truncated thoughts are also over-confident, so the equal mix does not damp them.
Outside Knowledge & Reasoning
Adaptive thinking with the same settings on five benchmarks of the other areas (run stopped after these):
| benchmark | metric | System 1 | adaptive | requests that thought |
|---|---|---|---|---|
| BFCL | case exact accuracy | 94.4 | 96.0 | 7 % |
| API-Bank | accuracy | 84.3 | 86.0 | 18 % |
| CLINC150+OOS (5,456 of 5,500) | macro-F1 | 93.3 | 94.4 | 10 % |
| ToolRet | nDCG@10 | 66.8 | 66.9 | 84 % |
| BANKING77 | macro-F1 | 88.0 | 85.0 | 19 % |
Small gains on tool calls and intents, none on retrieval, and a loss on BANKING77, whose training split System 1 was trained on: the untrained System 2 often overrules a correct System 1 with an over-confident wrong answer. Use thinking for reasoning questions; keep it off for classification, retrieval and tool routing.
Latency
Thinking costs time. Over the area's ten benchmarks, 66 % of the requests thought at least once. Thinking length:
median 8,192 tokens (the budget) on GPQA, HLE and ChessBench, about 5,400 on MMLU-Pro, 3,300 on BBH and 900 on GSM8K.
Estimated single-request latency on one idle B200 (System 1 ≈ 45 ms, thinking ≈ 250 tokens per second): about 0.05 s
without thinking and up to about 33 s with an 8,192-token thought; median 13.4 s over the area (90th percentile 33 s).
With speculative decoding (below) thinking runs at about 438 tokens per second: median 7.7 s, 90th percentile 19 s.
Lower threshold or think_budget to trade accuracy for speed; thinking: "off" keeps every decision in one pass.
MMLU-Pro and BBH (except six questions) ran with speculative decoding, which does not change the output distribution.
Six BBH questions with 18 options hit a vLLM error in the speculative-decoding read-out path and were answered by the same
server without speculative decoding; the current serve_decide.py reads answers through a path that works with
speculative decoding.
Images
The checkpoint contains Gemma-4's vision encoder (no audio encoder), so System 1 and System 2 both accept images. The decision head was trained on text; decisions over images are zero-shot.
| check | result |
|---|---|
| synthetic images: colour (8 options), shape (4), printed number (8), "is there a red object?" (yes/no) | 100 % on each (30 images each) |
| VL-RewardBench, 1,247 pairs, both presentation orders averaged | 78.4 % overall (general 55.8, hallucination 84.9, reasoning 76.0; macro 72.2) |
For reference, autotrust/JEV-27B-VL scores 78.3 % on VL-RewardBench with the same protocol.
Context length
The backbone's native context is 262,144 tokens (256K). We tested decisions that hinge on a single sentence placed at a random depth in long real text (concatenated PubMedQA abstracts): a yes/no question and a 16-option question, 10 of each per length, on one B200 with vLLM, one request at a time.
| prompt length | yes/no correct | 16-option correct | mean probability on the right answer | median latency |
|---|---|---|---|---|
| 4K | 10/10 | 10/10 | 0.999 | 0.15 s |
| 32K | 10/10 | 10/10 | 0.999 | 1.6 s |
| 64K | 10/10 | 10/10 | 0.999 | 5.0 s |
| 128K | 10/10 | 10/10 | 1.000 | 17.8 s |
Lengths above 128K have not been tested yet.
Many options
choice takes 2–256 options. Up to 16 are read in one pass with the trained labels A–P. More options are read in groups
of at most 16 (in parallel), then a final of 16; every option is read and none is pruned (strategy: "tournament", the
default). Zero-shot intent classification, all options offered at once, 400 test utterances per row:
| test | options | accuracy |
|---|---|---|
| MASSIVE (en) | 59 | 91.2 % |
| BANKING77 | 77 | 81.5 % |
| CLINC150 | 150 | 95.5 % |
| CLINC150 utterances among CLINC150 + BANKING77 + MASSIVE intents | 255 | 89.0 % |
| BANKING77 utterances among the same 255 intents | 255 | 75.2 % |
The 255-option sets merge three catalogues with overlapping intents, so part of the drop comes from near-duplicate labels.
A single pass with labels beyond P (strategy: "single") is about 3× faster but less accurate here (CLINC150: 88.5 %
against 95.2 %), so it is not the default. The BANKING77 and CLINC150 training splits are part of the training data (see
below).
Quick start (vLLM)
hf download autotrust/GEV-26B-Decide --local-dir GEV-26B-Decide
bash GEV-26B-Decide/serve.sh # vLLM on :8000; one GPU with 80 GB or more
serve.sh runs serve_decide.py: the standard vLLM OpenAI server (same flags as vllm serve) with a POST /v1/decide
route. It loads the backbone once: plain requests are System 2, and requests for the LoRA module jev-decision
(adapter_vllm/: backbone LoRA + the decision head as an lm_head LoRA) are System 1. It needs a vLLM build with
Gemma-4 support plus patches/vllm-gemma4-lm-head-lora.patch (LoRA on Gemma-4's tied lm_head, vocabulary 262,144);
tested with a vLLM development build from September 2026.
System 1 and adaptive thinking: POST /v1/decide
curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
"kind": "choice",
"state": "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball.",
"question": "How much does the ball cost?",
"options": ["$0.10", "$0.05", "$1.00", "$0.55"],
"thinking": "auto"}'
| field | value |
|---|---|
kind |
noul: yes/no, probabilities for ["false", "true"] · score: 0–5 · choice: your options |
state |
what the decision is about: a string, a JSON object, or a list mixing text and images ["Photo: ", {"image": "https://… or data:…"}] |
question |
one question about the state |
options |
choice only: 2–256 strings |
thinking |
"off" (default: System 1 only), "auto" (adaptive), "on" (always think); noul and choice. Switch on "auto" for reasoning questions; keep "off" for classification, retrieval and tool routing |
threshold |
System 1 confidence below which "auto" thinks (default 0.8) |
think_budget |
maximum thinking tokens; default: no cap beyond the context window |
chat_template_kwargs |
passed to the base model's chat template, as in its chat API (Gemma-4 has thinking on/off only, so there is no reasoning_effort setting) |
strategy |
more than 16 options: "tournament" (default), "single", "permute" |
return_reasoning / debug |
include System 2's reasoning / the System 1 and System 2 distributions |
The response has options, probabilities, choice, choice_index, usage and, when thinking was requested,
thinking: {"used": true, "think_tokens": …, "think_seconds": …, "finished_within_budget": …}. GET /v1/decide/info
lists the defaults.
import requests
def decide(kind, state, question, options=None, thinking="auto"):
body = {"kind": kind, "state": state, "question": question, "thinking": thinking, **({"options": options} if options else {})}
r = requests.post("http://localhost:8000/v1/decide", json=body).json()
return dict(zip(r["options"], r["probabilities"])), r.get("thinking", {}).get("used")
decide("noul", "John was born on 29 February 1996.", "Was John's 7th birthday celebrated on a 29 February?")
System 2
requests.post("http://localhost:8000/v1/chat/completions", json={
"model": "autotrust/GEV-26B-Decide",
"messages": [{"role": "user", "content": "In one sentence, what is safety stock?"}],
"max_tokens": 200, "chat_template_kwargs": {"enable_thinking": False}})
Speed (one B200, vLLM)
- System 1: median 45 ms for a single request; 257 decisions per second with 64 concurrent clients. vLLM matches the
transformersengine below to a mean largest probability difference of 0.015 (300 held-out decisions). - Adaptive thinking adds nothing when it does not think, and about 1 second per 250 thinking tokens when it does (see Latency).
Faster thinking with speculative decoding. MTP=1 bash GEV-26B-Decide/serve.sh adds Google's 0.9 GB draft model for
this backbone (--speculative-config '{"model": "google/gemma-4-26B-A4B-it-assistant", "num_speculative_tokens": 4}').
On one B200 it speeds up System 2 about 1.8–1.9×: 247 → 438 tokens per second for a single request, 7,072 → 13,505 tokens
per second at 128 concurrent requests (mean acceptance length 3.5–3.7 of 4). System 1 decisions are unchanged, but System 1
throughput at high concurrency drops (257 → 140 decisions per second with 64 clients); use it when you mostly think.
serve_decide.py reads answers with the token restriction that works under speculative decoding, which needs
--max-logprobs 256 (set in serve.sh).
If you call /v1/completions for System 1 yourself, pass top_k: 0 and top_p: 1.0: the model's generation config sets
top_k=64 and top_p=0.95, which vLLM applies as request defaults and which would truncate the returned probabilities.
Usage (transformers + peft, System 1)
import json, torch
from huggingface_hub import snapshot_download
from peft import PeftModel
from safetensors.torch import load_file
from transformers import AutoTokenizer, Gemma4ForConditionalGeneration
d = snapshot_download("autotrust/GEV-26B-Decide")
tok = AutoTokenizer.from_pretrained(d)
base = Gemma4ForConditionalGeneration.from_pretrained(d, dtype=torch.bfloat16, device_map="cuda")
# System 2: `base` is gemma-4-26B-A4B-it unchanged; use base.generate(...) (text or images).
# System 1: adapter merged in memory + head
m = PeftModel.from_pretrained(base, f"{d}/adapter").merge_and_unload().eval()
backbone = m.model
jc, T = json.load(open(f"{d}/judge_config.json")), json.load(open(f"{d}/calibration.json"))["per_kind"]
head = load_file(f"{d}/head.safetensors"); W, b = head["proj.weight"].cuda(), head["proj.bias"].cuda()
@torch.no_grad()
def decide(kind, state, question, options):
lines = options if kind != "choice" else [f"{'ABCDEFGHIJKLMNOP'[i]}) {o}" for i, o in enumerate(options)]
text = f"[kind] {kind}\n[state] {state}\n[question] {question}\n[options]\n" + "\n".join(lines) + "\n[decision]:"
ids = torch.tensor([[tok.bos_token_id] + tok.encode(text, add_special_tokens=False)], device="cuda")
h = backbone(input_ids=ids, use_cache=False).last_hidden_state[0, -1].float()
z = 30.0 * torch.tanh((W @ h + b) / 30.0)
s, _ = jc["slots"]["ranges"][kind]
return dict(zip(options, torch.softmax(z[s:s + len(options)] / T[kind], 0).tolist()))
print(decide("noul", "Customer says the parcel arrived damaged and wants their money back.",
"Is the customer asking for a refund?", ["false", "true"]))
This path reads up to 16 options per pass; for more, use the server (or read groups of 16 and a final, as above).
Engine and read-out
- Template
bare-v1, prefixed with<bos>:[kind] … [state] … [question] … [options] A) … [decision]: - Read-out: the final-norm hidden state of the last token goes through a linear fp32 head (hidden 2,816 → 24 slots), soft-capped at 30 like Gemma's own logits; inactive slots are masked, the logits are divided by the per-kind temperature and softmaxed. Nothing is generated.
Temperatures
| table | noul | choice | score |
|---|---|---|---|
calibration.json (default; used in the Decision Index run) |
1.003 | 1.017 | 0.999 |
calibration_gold.json (calibrated against ground-truth answers) |
1.214 | 1.098 | 1.000 |
Use calibration_gold.json when you gate automatic actions on confidence.
Training data (disclosure)
System 1 was trained on teacher distributions and ground-truth decision data. The ground-truth data includes the public training splits of some datasets whose test splits the Decision Index uses (among them BANKING77 and CLINC150); no test split of any benchmark was used, and suite items were excluded before training. The list has been provided to the Decision Index maintainers. MMMU / MMMU-Pro are not valid evaluations for this model. The adaptive-thinking settings were chosen on data outside the Decision Index suite.
Limitations
- Where the teacher is wrong, System 1 often is too, and it can be confidently wrong: on HLE the questions it answers without thinking are 1.6 % correct. Thinking does not help classification or retrieval and can hurt tasks System 1 was trained on (BANKING77 macro-F1 88.0 → 85.0). Adaptive thinking helps most on science, code, math and logic, does not help on chess or expert-level HLE questions, and can slightly lower accuracy on commonsense questions (CommonsenseQA 86.7 → 85.0).
- Thinking is slow on hard inputs: on expert-level and chess questions it often uses the full 8,192-token budget, and the adaptive mode does not meet the Decision Index latency limit on the Knowledge & Reasoning benchmarks.
- The Decision Index score above uses adaptive thinking on the Knowledge & Reasoning area and System 1 on the other four areas; it is our own scoring, not a board result.
- Weaker than JEV-27B on long structured inputs (e.g. the Decision Index's Home appliance simulator and POP909).
- HLE: below chance with System 1, like every open entry on the board.
- Decisions over images are zero-shot; contexts above 128K tokens are untested.
- The one-engine vLLM setup needs the bundled vLLM patch;
noulandscoreaccept only their canonical options. - English-centric; not for high-stakes decisions without confidence gating.
Files
model-*.safetensors · config.json · processor_config.json · tokenizer* · chat_template.jinja · generation_config.json
google/gemma-4-26B-A4B-it, unchanged (System 2; text + image input)
adapter/ System 1 LoRA (peft), for the transformers path
head.safetensors 24-slot decision head (fp32): proj.weight [24, 2816], proj.bias [24]
judge_config.json slot layout, verbalizer ids, softcap, read-out
calibration.json per-kind temperatures (default)
calibration_gold.json per-kind temperatures calibrated against ground-truth answers
adapter_vllm/ System 1 for vLLM: backbone LoRA + the head as an lm_head LoRA, plus decision_head.json
serve_decide.py · serve.sh
vLLM server with POST /v1/decide (System 1, adaptive thinking) next to the OpenAI endpoints
patches/ vLLM patch: LoRA on Gemma-4's tied lm_head
reports/ adaptive thinking: validation summary, Decision Index recomputation, latency summary, other-area sample, games
videos/ adaptive-thinking videos (Minesweeper, Connect Four, Wordle, Sudoku)
License
Apache-2.0 for the adapter, head and calibration files; base model under the Gemma 4 terms (https://ai.google.dev/gemma/docs/gemma_4_license). Not affiliated with TypeSafe AI.
- Downloads last month
- -
Model tree for autotrust/GEV-26B-Decide
Evaluation results
- Decision Index (balanced skill) on Decision Index suite 0.2 (edition 0.2.1)self-reported62.480