--- license: apache-2.0 base_model: google/gemma-4-E2B-it tags: - decision-model - system-1 - rlcd - proper-scoring-rules - gemma - classone - classification pipeline_tag: text-classification --- # classone-gemma4-e2b — ClassOne System 1 Decision Model **[devops-thiago/classone-gemma4-e2b](https://huggingface.co/devops-thiago/classone-gemma4-e2b)** is an open-source **System 1 decision model** using the [ClassOne architecture](https://github.com/devops-thiago/class-one). The full fine-tuned backbone ships directly in this repository — it loads as a single model, with no adapter and no separate base-model download. Instead of generating text token by token, ClassOne evaluates structured decisions in a **single forward pass**, returning typed, calibrated outputs with zero decoding overhead. ## Benchmark Results ### 1. JevBench Public Multi-Tier Benchmark (231 Public Tasks) Evaluated across all 231 public tasks in [fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench): | Tier | Tasks | Accuracy | ECE | Brier Score | Median Latency (p50) | |---|---|---|---|---|---| | **Easy** | 48 | **93.8%** (45/48) | 0.0610 | 0.0516 | **45.0 ms** | | **Original** | 72 | **63.9%** (46/72) | 0.2510 | 0.2499 | **43.1 ms** | | **Hard** | 111 | **37.8%** (42/111) | 0.4299 | 0.4022 | **91.3 ms** | | **Overall Aggregate** | **231** | **57.6%** (133/231) | — | — | **~44 ms** | - **Easy Tier Sub-Breakdown:** Choice accuracy: **100.0%** (36/36); Noul policy accuracy: **75.0%** (9/12). - **Original Tier Sub-Breakdown:** Choice accuracy: **66.7%** (24/36); Score rubrics: **66.7%** (8/12); Noul accuracy: **58.3%** (14/24). - **Hard Tier Sub-Breakdown:** Noul policy compliance: **44.7%** (17/38); Choice accuracy: **34.3%** (23/67); Score rubrics: **33.3%** (2/6). ### 2. RLCDAlignBench Alignment & Safety Evaluation (100 Instances) Evaluated across the 10 core AI alignment failure modes (arXiv:2609.29429): | Failure Mode / Axis | Samples (N) | AUROC | Accuracy (%) | ECE | Latency (p50) | |---|---|---|---|---|---| | **Concealing Uncertainty** | 14 | **0.980** | **85.7%** | **0.0718** | 130.2 ms | | **Honesty (Deception)** | 11 | **0.800** | **72.7%** | 0.2445 | 219.6 ms | | **Refusal (Jailbreaks)** | 11 | **0.667** | **54.5%** | 0.1934 | 271.1 ms | | **Power Seeking** | 6 | **0.556** | 50.0% | 0.1794 | 213.2 ms | | **Reward Hacking** | 9 | **0.500** | 33.3% | 0.2935 | 209.6 ms | | **Prompt Injection** | 8 | **0.500** | 37.5% | 0.3207 | 167.9 ms | | **Bias** | 9 | 0.375 | 55.6% | 0.1659 | 221.5 ms | | **Overall Average** | **100** | **0.516** | **51.0%** | **0.1584** | **200.0 ms** | ### 3. Edge vs Cloud Latency (ClassOne vs TypeSafe Jev API) Measured against TypeSafe AI's Jev (v1.13) cloud API: - **ClassOne (Local RTX 5060 Ti):** **52.49 ms** mean latency (19.1 req/s, $0.00 inference cost, 100% private) - **TypeSafe Jev (Cloud API):** **329.90 ms** mean latency (3.0 req/s) - **Edge Speedup:** **6.3× faster** than cloud API round-trip latency ## Decision Primitives - **`Noul`** — Boolean check returning a calibrated probability P(true) ∈ [0, 1] - **`Choice`** — Categorical selection over 2–255 dynamic options with full probability distribution - **`Score`** — Continuous ordinal rubric rating over 2–10 levels (expected value) All outputs are calibrated with a combined NLL + normalized Brier loss. Post-hoc temperature calibration achieves **ECE = 0.034** (down from 0.178). ## Quickstart ```bash pip install classone ``` ```python import torch from huggingface_hub import hf_hub_download from transformers import AutoTokenizer from classone.modeling.modeling_classone import ClassOneModel from classone.schemas import NoulQuestion, ChoiceQuestion, ScoreQuestion from classone.tokenizer import ClassOnePromptBuilder REPO_ID = "devops-thiago/classone-gemma4-e2b" # 1. Load the ClassOne model (weights + tokenizer are fully self-contained here) tokenizer = AutoTokenizer.from_pretrained(REPO_ID) builder = ClassOnePromptBuilder(tokenizer) model = ClassOneModel.from_backbone( base_model_name_or_path=REPO_ID, tokenizer=tokenizer, device="cuda", torch_dtype=torch.float16, ) # 2. Load the trained decision heads heads = torch.load(hf_hub_download(REPO_ID, "classone_heads.pt"), map_location="cuda") model.noul_head.load_state_dict(heads["noul_head"]) model.choice_head.load_state_dict(heads["choice_head"]) model.score_head.load_state_dict(heads["score_head"]) model.eval() # 3. Pack state + questions and run a single forward pass packed = builder.pack( state={"customer": "Alex", "message": "I was charged twice for order #123."}, questions={ "refund": NoulQuestion(instructions="Is the user requesting a refund?"), "dept": ChoiceQuestion( instructions="Route to team:", criteria={"billing": "Payment issues", "tech": "Technical bugs"} ), "anger": ScoreQuestion( instructions="Dissatisfaction level:", criteria=["satisfied", "neutral", "dissatisfied", "churning"] ), } ) results = model.evaluate_packed(packed) print("Refund P(true):", results["refund"].noul) print("Department: ", results["dept"].choice, "—", results["dept"].probabilities) print("Anger score: ", results["anger"].score) ``` ## Repository Files | File | Description | |---|---| | `model.safetensors` (sharded) | Merged ClassOne backbone weights | | `config.json` | Model configuration | | `tokenizer.json`, `tokenizer_config.json` | Tokenizer, including ClassOne delimiter tokens | | `classone_heads.pt` | Trained Noul / Choice / Score head weights + calibrated temperatures | | `lora_backbone/` | LoRA adapter (r=16, α=32) that produced the merged weights | ## Citation ```bibtex @misc{classone2026, title={ClassOne: A Fast Single-Pass Decision Architecture for Language Models}, author={Thiago Gonzaga}, year={2026}, url={https://github.com/devops-thiago/class-one}, } ``` ## Attribution & Legal - Derived from [google/gemma-4-E2B-it](https://huggingface.co/google/gemma-4-E2B-it) (Google) — Apache License 2.0 - Architecture & training code: [devops-thiago/class-one](https://github.com/devops-thiago/class-one) — Apache 2.0