System One Model (using Gemma 4 12B)

Using encoder and decoder from google/gemma-4-12B with its 262k-token language-model head replaced by a low dimentional vector that outputs one logit per answer. The forward pass gives the 26 logits. Questions (noul, choice, score) are answered by softmax over the logits for each question.

Benchmark on RTX 5090

Median latency is 34.1ms for a warm single-question decision.

Warm single-question decisions (latency_bench, 55 calls):
    p50  34.1ms
    p95  35.0ms
    max  36.3ms

Prefix cache (8 endings):
    decisions changed:    0 of 8
    batched vs unbatched: 64.3ms vs 270ms

5-question request: 166ms

Comparison: Strands Decider 2B and the hosted Jev API

Same 39 labelled cases (15 noul, 12 choice, 12 score) sent through /v1/systemone schema to each system. Local models ran on an Apple Silicon Mac (64 GB, MPS, bf16)

metric this model (Gemma 4 12B, local) Strands Decider 2B v19 (local) TypeSafe Jev API (hosted)
Mario 5/5 5/5 5/5
Tetris 40/40 40/40 40/40
noul accuracy 15/15 15/15 15/15
noul Brier (lower is better) 0.000003 0.0338 0.00035
choice accuracy 11/12 12/12 12/12
score accuracy 11/12 5/12 12/12
score mean error (levels) 0.09 0.48 0.004
1-question p50 latency 117 ms 96 ms 144 ms
3-question p50 latency 332 ms 121 ms 164 ms
Size ~24 GB ~2B params n/a

Test suite: 141 passed. Small hand-written sample

Doom played by the model.

Tetris played by the model.

Super Mario Bros. World 1-1 played by the model.

Flappy Bird played by the model.

Model API

Requires transformers>=5.17. The repo ships custom code, so pass trust_remote_code=True.

from transformers import AutoModelForMultimodalLM, AutoTokenizer

repo = "stephenlb/system-one-model"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMultimodalLM.from_pretrained(
    repo, trust_remote_code=True, dtype="bfloat16", device_map="auto"
)

result = model.system_one(
    tokenizer,
    state="I was charged twice for order A-104.",
    questions={
        "refund": {"type": "noul", "instructions": "Does the text request a refund?"},
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this?",
            "criteria": {"billing": "Charges", "returns": "Refunds"},
        },
    },
)
print(result["answers"])

Pipeline API

from transformers import pipeline

pipe = pipeline("system-one", model=repo, trust_remote_code=True, dtype="bfloat16", device_map="auto")
print(pipe({"state": "I was charged twice.", "questions": {...}}))

Output

{"answers": {
  "refund": {"type": "noul", "noul": 0.93},
  "team": {"type": "choice", "choice": "billing",
           "probabilities": {"billing": 0.97, "returns": 0.03}, "confidence": 0.8},
  "urgency": {"type": "score", "score": 3.1, "legend": {"0": "..."},
              "probabilities": {"0": 0.0, "1": 0.1}, "confidence": 0.5}}}

noul is the probability of yes. choice returns the top label with probabilities per option. score returns the expected level over an ordered criteria list. Pass temperature= (default 0.7) to sharpen or soften.

Examples

Outputs below are real results from this model (pipeline defaults, temperature 0.7).

noul: yes/no probability

noul is the probability the answer is yes. Ask several yes/no questions about one text in a single call.

pipe({
    "state": "Please cancel my subscription and refund this month's charge.",
    "questions": {
        "refund": {"type": "noul", "instructions": "Does the text request a refund?"},
        "cancel": {"type": "noul", "instructions": "Does the text ask to cancel a subscription?"},
        "sports": {"type": "noul", "instructions": "Does the text talk about sports?"},
    },
})["answers"]
# {"refund": {"type": "noul", "noul": 0.9996},
#  "cancel": {"type": "noul", "noul": 0.9998},
#  "sports": {"type": "noul", "noul": 0.0023}}

Optional criteria overrides the yes/no wording: {"true": "...", "false": "..."}.

Another text, different questions: tone and content checks.

pipe({
    "state": "Thanks so much for the quick help, the new dashboard looks great!",
    "questions": {
        "polite": {"type": "noul", "instructions": "Is the text polite?"},
        "complaint": {"type": "noul", "instructions": "Does the text contain a complaint?"},
        "question": {"type": "noul", "instructions": "Does the text ask a question?"},
    },
})["answers"]
# {"polite":    {"type": "noul", "noul": 0.9991},
#  "complaint": {"type": "noul", "noul": 0.0009},
#  "question":  {"type": "noul", "noul": 0.0080}}

choice: pick one option

criteria maps each option label to a description. You get the top label, a probability for every option, and a confidence (1.0 means all mass on one option, 0.0 means uniform). Up to 26 options.

pipe({
    "state": "My package says delivered but nothing arrived at my door.",
    "questions": {
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this?",
            "criteria": {
                "billing": "Charges, invoices, payment problems",
                "shipping": "Delivery status, delays, lost packages",
                "returns": "Exchanges, refunds, wrong or damaged items",
            },
        },
    },
})["answers"]
# {"team": {"type": "choice", "choice": "shipping",
#           "probabilities": {"shipping": 0.9977, "returns": 0.0016, "billing": 0.0007},
#           "confidence": 0.984}}

Intent classification with four options:

pipe({
    "state": "Can you add a dark mode to the mobile app? My eyes hurt at night.",
    "questions": {
        "intent": {
            "type": "choice",
            "instructions": "What is the intent of the message?",
            "criteria": {
                "bug_report": "Something is broken or not working as expected",
                "feature_request": "Asks for new functionality or an improvement",
                "praise": "Compliments or thanks",
                "question": "Asks how to do something",
            },
        },
    },
})["answers"]
# {"intent": {"type": "choice", "choice": "feature_request",
#             "probabilities": {"feature_request": 0.9781, "bug_report": 0.0094,
#                               "question": 0.0079, "praise": 0.0046},
#             "confidence": 0.907}}

score: ordered scale

criteria is an ordered list of levels (lowest first, 2 to 26 levels). score is the expected level sum(i * p_i), so 1.94 below means mostly level 2 with a little level 1. legend maps each level to its description.

pipe({
    "state": "The checkout page crashes every time I press Pay and I can't buy anything.",
    "questions": {
        "severity": {
            "type": "score",
            "instructions": "How severe is the reported issue?",
            "criteria": [
                "Cosmetic; no impact to functionality",
                "Broken or degraded feature, but workaround exists",
                "Blocking issue; no workaround exists",
            ],
        },
    },
})["answers"]["severity"]
# {"type": "score", "score": 1.942,
#  "legend": {"0": "Cosmetic; ...", "1": "Broken or degraded ...", "2": "Blocking issue; ..."},
#  "probabilities": {"0": 0.0018, "1": 0.0542, "2": 0.9440}, "confidence": 0.796}

A five-point scale on mixed feedback. The distribution is spread out, so score lands between levels and confidence is low:

pipe({
    "state": "The delivery was two days late, the box was dented, but the product itself works fine.",
    "questions": {
        "satisfaction": {
            "type": "score",
            "instructions": "How satisfied is the customer overall?",
            "criteria": ["Very dissatisfied", "Dissatisfied", "Neutral", "Satisfied", "Very satisfied"],
        },
    },
})["answers"]["satisfaction"]
# {"type": "score", "score": 1.408,
#  "legend": {"0": "Very dissatisfied", "1": "Dissatisfied", "2": "Neutral",
#             "3": "Satisfied", "4": "Very satisfied"},
#  "probabilities": {"0": 0.1144, "1": 0.5705, "2": 0.1635, "3": 0.0957, "4": 0.0560},
#  "confidence": 0.223}

Mixing types in one call

All questions in a call see the same state and run in one batched forward pass. state may also be a dict or list; it is serialized to JSON.

pipe({
    "state": {"customer": "A-104", "message": "The checkout page crashes when I press Pay."},
    "questions": {
        "urgent": {"type": "noul", "instructions": "Does the text express urgency?"},
        "sentiment": {
            "type": "choice",
            "instructions": "What is the sentiment of the text?",
            "criteria": {
                "positive": "Happy or satisfied",
                "neutral": "Factual, no strong emotion",
                "negative": "Unhappy or frustrated",
            },
        },
        "severity": {
            "type": "score",
            "instructions": "How severe is the reported issue?",
            "criteria": ["Cosmetic", "Degraded, workaround exists", "Blocking"],
        },
    },
}, temperature=0.5, include_letter_logits=True)

temperature below 0.7 sharpens the probabilities; above it, softens them. It never changes which option wins. include_letter_logits=True adds the raw A-Z logits per question under letter_logits.

Raw logits

enc = tokenizer(prompts, return_tensors="pt", padding=True)  # padding_side="left"
logits = model.letter_logits(enc.input_ids, enc.attention_mask)  # [batch, 26]

Blocks.ai

We needed an open weight model that offered the capabilities of Jev System One Model. Most common AI Agents require decisions making. The System One model approach is a great new way to do this. Blocks.ai is the secure network for the Internet of Agents (IoA). Whether connecting agents to users in your organization or making your agents available for public use, Blocks.ai makes your agents securely discoverable and callable by everyone who needs them most. We Rebuilt Jev's API on an Open Model and Used It to Play Doom. TypeSafe AI's Jev turns unstructured input into typed decisions and probabilities. We rebuilt the API with an open model, reached 113ms median latency on an M4 Mac, and used it to play Doom.

Github Source System One Model

https://github.com/stephenlb/system-one-model

Downloads last month
45
Safetensors
Model size
12B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for stephenlb/system-one-model

Finetuned
(70)
this model