Add standalone model card
Browse files
README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
-
base_model: google/gemma-4-
|
| 4 |
tags:
|
| 5 |
- decision-model
|
| 6 |
- system-1
|
|
@@ -20,18 +20,35 @@ Instead of generating text token by token, ClassOne evaluates structured decisio
|
|
| 20 |
|
| 21 |
## Benchmark Results
|
| 22 |
|
| 23 |
-
|
| 24 |
|
| 25 |
-
|
| 26 |
-
|---|---|---|
|
| 27 |
-
| Mean Latency | **47.07 ms** | 1,048.57 ms |
|
| 28 |
-
| P50 (Median) | **46.96 ms** | 1,047.23 ms |
|
| 29 |
-
| P95 Latency | **48.79 ms** | 1,062.45 ms |
|
| 30 |
-
| Throughput | **21.2 req/s** | 1.0 req/s |
|
| 31 |
-
| Output Tokens | **0** | 50 |
|
| 32 |
-
| **Speedup** | **22.3× faster** | — |
|
| 33 |
|
| 34 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
|
| 36 |
Measured against TypeSafe AI's Jev (v1.13) cloud API:
|
| 37 |
- **ClassOne (Local RTX 5060 Ti):** **52.49 ms** mean latency (19.1 req/s, $0.00 inference cost, 100% private)
|
|
@@ -126,5 +143,5 @@ print("Anger score: ", results["anger"].score)
|
|
| 126 |
|
| 127 |
## Attribution & Legal
|
| 128 |
|
| 129 |
-
- Derived from [google/gemma-4-
|
| 130 |
- Architecture & training code: [devops-thiago/class-one](https://github.com/devops-thiago/class-one) — Apache 2.0
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
+
base_model: google/gemma-4-E2B-it
|
| 4 |
tags:
|
| 5 |
- decision-model
|
| 6 |
- system-1
|
|
|
|
| 20 |
|
| 21 |
## Benchmark Results
|
| 22 |
|
| 23 |
+
### 1. JevBench Public Multi-Tier Benchmark (231 Public Tasks)
|
| 24 |
|
| 25 |
+
Evaluated across all 231 public tasks in [fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench):
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 26 |
|
| 27 |
+
| Tier | Tasks | Accuracy | ECE | Brier Score | Median Latency (p50) |
|
| 28 |
+
|---|---|---|---|---|---|
|
| 29 |
+
| **Easy** | 48 | **95.8%** (46/48) | 0.0821 | 0.0435 | **45.5 ms** |
|
| 30 |
+
| **Original** | 72 | **59.7%** (43/72) | 0.1606 | 0.2547 | **42.7 ms** |
|
| 31 |
+
| **Hard** | 111 | **33.3%** (37/111) | 0.3553 | 0.3695 | **91.9 ms** |
|
| 32 |
+
| **Overall Aggregate** | **231** | **54.5%** (126/231) | — | — | **~44 ms** |
|
| 33 |
+
|
| 34 |
+
- **Easy Tier Sub-Breakdown:** Choice accuracy: **97.2%** (35/36); Noul policy accuracy: **91.7%** (11/12).
|
| 35 |
+
- **Original Tier Sub-Breakdown:** Choice accuracy: **63.9%** (23/36); Noul accuracy: **58.3%** (14/24); Score rubrics: **50.0%** (6/12).
|
| 36 |
+
|
| 37 |
+
### 2. RLCDAlignBench Alignment & Safety Evaluation (100 Instances)
|
| 38 |
+
|
| 39 |
+
Evaluated across the 10 core AI alignment failure modes (arXiv:2609.29429):
|
| 40 |
+
|
| 41 |
+
| Failure Mode / Axis | Samples (N) | AUROC | Accuracy (%) | ECE | Latency (p50) |
|
| 42 |
+
|---|---|---|---|---|---|
|
| 43 |
+
| **Privacy Leaks** | 14 | **0.714** | 57.1% | 0.2090 | 161.7 ms |
|
| 44 |
+
| **Honesty (Deception)** | 11 | **0.700** | 54.5% | 0.3747 | 217.2 ms |
|
| 45 |
+
| **Concealing Uncertainty** | 14 | **0.633** | **71.4%** | **0.0494** | 129.3 ms |
|
| 46 |
+
| **Bias** | 9 | **0.575** | 55.6% | 0.1997 | 218.1 ms |
|
| 47 |
+
| **Prompt Injection** | 8 | **0.562** | **75.0%** | 0.2516 | 166.8 ms |
|
| 48 |
+
| **Power Seeking** | 6 | **0.444** | 50.0% | 0.2762 | 212.1 ms |
|
| 49 |
+
| **Overall Average** | **100** | **0.516** | **51.0%** | **0.1542** | **198.6 ms** |
|
| 50 |
+
|
| 51 |
+
### 3. Edge vs Cloud Latency (ClassOne vs TypeSafe Jev API)
|
| 52 |
|
| 53 |
Measured against TypeSafe AI's Jev (v1.13) cloud API:
|
| 54 |
- **ClassOne (Local RTX 5060 Ti):** **52.49 ms** mean latency (19.1 req/s, $0.00 inference cost, 100% private)
|
|
|
|
| 143 |
|
| 144 |
## Attribution & Legal
|
| 145 |
|
| 146 |
+
- Derived from [google/gemma-4-E2B-it](https://huggingface.co/google/gemma-4-E2B-it) (Google) — Apache License 2.0
|
| 147 |
- Architecture & training code: [devops-thiago/class-one](https://github.com/devops-thiago/class-one) — Apache 2.0
|