Zero-Shot Classification
PEFT
Safetensors
English
lev
system-one
decision-model
calibrated-decisions
classification
routing
moderation
lora
Instructions to use interfaze-ai/lev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use interfaze-ai/lev with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B") model = PeftModel.from_pretrained(base_model, "interfaze-ai/lev") - Notebooks
- Google Colab
- Kaggle
docs: report lev vs Jev on all 13 S1Bench subsets
Browse files- .gitattributes +1 -0
- README.md +36 -41
- assets/accuracy-by-subset.png +0 -0
- assets/accuracy-per-parameter.png +2 -2
- assets/compute.png +0 -0
- assets/latency.png +0 -0
- assets/leaderboard.png +2 -2
.gitattributes
CHANGED
|
@@ -37,3 +37,4 @@ assets/leaderboard.png filter=lfs diff=lfs merge=lfs -text
|
|
| 37 |
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
| 38 |
assets/accuracy-per-parameter.png filter=lfs diff=lfs merge=lfs -text
|
| 39 |
assets/hero.png filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 37 |
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
| 38 |
assets/accuracy-per-parameter.png filter=lfs diff=lfs merge=lfs -text
|
| 39 |
assets/hero.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
+
assets/accuracy-by-subset.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -26,18 +26,16 @@ lev answers typed questions about a piece of context in a single forward pass. Y
|
|
| 26 |
|
| 27 |
<div align="center" style="line-height: 1;"><img src="https://img.shields.io/badge/license-Apache--2.0-2a78d6?style=flat-square" alt="Apache-2.0" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/base-Qwen3.5--4B-2a78d6?style=flat-square" alt="Qwen3.5-4B" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/output%20tokens-0-2a78d6?style=flat-square" alt="Zero output tokens" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/API-%2Fv1%2Fsystemone-2a78d6?style=flat-square" alt="/v1/systemone compatible" style="display: inline-block; vertical-align: middle; margin: 2px;"> <a href="https://github.com/Abhinavexists/lev"><img src="https://img.shields.io/badge/code-GitHub-14181f?style=flat-square&logo=github" alt="GitHub" style="display: inline-block; vertical-align: middle; margin: 2px;"></a></div>
|
| 28 |
|
| 29 |
-
<h2 align="center">
|
| 30 |
|
| 31 |
-
<p align="center"><strong>Qwen3.5-4B + LoRA
|
| 32 |
|
| 33 |
-
<p align="center"><a href="#quickstart"><strong>Quickstart</strong></a> · <a href="#self-hosting-a-jev-compatible-http-server">Self-hosting</a> · <a href="#benchmarks">Benchmarks</a> · <a href="#why-it-works">Why it works</a> · <a href="#training">Training</a> · <a href="https://github.com/Abhinavexists/lev">GitHub</a></p>
|
| 34 |
|
| 35 |
-
**No generated tokens
|
| 36 |
-
|
| 37 |
-
**lev cannot return a label outside your options.** The answer space is the option set you send, so every response is well-formed by construction. lev can still pick the wrong option: this is a structural guarantee, not a guarantee of correctness.
|
| 38 |
|
| 39 |
| Question | You give | You get |
|
| 40 |
-
|---|---|---|
|
| 41 |
| `noul` | a yes/no question | `noul` = p(yes) |
|
| 42 |
| `choice` | instructions + options (name → description or `null`) | `choice`, `probabilities`, `confidence` |
|
| 43 |
| `score` | instructions + 2–10 ordered levels | `score` (expected level), `probabilities`, `confidence` |
|
|
@@ -133,42 +131,42 @@ print(response.answers["team"].choice, response.answers["bug"].noul) # technica
|
|
| 133 |
|
| 134 |
### S1Bench
|
| 135 |
|
| 136 |
-
Six S1Bench subsets, 1,999 items. lev and TypeSafe Jev were run through the same harness on the same task files. Accuracy, best in each row in bold:
|
| 137 |
-
|
| 138 |
<p align="center"><img src="assets/accuracy-by-subset.png" alt="lev and Jev accuracy on each S1Bench subset" width="100%"></p>
|
| 139 |
|
| 140 |
-
| subset | task | **lev** | Jev |
|
| 141 |
-
|---|---|--:|--:|--:|
|
| 142 |
-
|
|
| 143 |
-
|
|
| 144 |
-
| massive-
|
| 145 |
-
|
|
| 146 |
-
|
|
| 147 |
-
|
|
| 148 |
-
|
|
| 149 |
-
|
| 150 |
-
|
| 151 |
-
|
| 152 |
-
|
| 153 |
-
|
| 154 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 155 |
|
| 156 |
<p align="center"><img src="assets/accuracy-per-parameter.png" alt="Macro accuracy against parameter count" width="100%"></p>
|
| 157 |
|
| 158 |
-
#### Where Jev leads
|
| 159 |
-
|
| 160 |
-
- **Minimal-edit pairs:** paws (−10.4 points) and vitaminc (−10.8), where two inputs differ by one swapped word or one changed number.
|
| 161 |
-
- **boolq and 60-class intent:** 3.0 and 2.3 points behind.
|
| 162 |
-
- **End-to-end latency from a laptop:** Jev's hosted API answered in 344–357 ms median, lev on one Modal H100 in 414–463 ms.
|
| 163 |
-
|
| 164 |
-
lev leads on safety moderation (+3.2) and helpfulness rating (+5.6).
|
| 165 |
-
|
| 166 |
### Held-out split
|
| 167 |
|
| 168 |
A held-out split of the 29 training sources, with no row shared with training:
|
| 169 |
|
| 170 |
| metric | value |
|
| 171 |
-
|---|---|
|
| 172 |
| weighted accuracy | 0.807\* |
|
| 173 |
| expected calibration error | 0.061\* (0.180 before calibration) |
|
| 174 |
| banking77 (77 intents) | 0.980 |
|
|
@@ -183,9 +181,7 @@ A held-out split of the 29 training sources, with no row shared with training:
|
|
| 183 |
|
| 184 |
A call is one batched forward pass over every question, so compute stays flat from one question to eight, and a 60-option choice costs the same as a yes/no.
|
| 185 |
|
| 186 |
-
|
| 187 |
-
|
| 188 |
-
Most of lev's round trip from a laptop is network and Modal's ingress; its compute is 69 ms. Self-hosted next to your application, that network hop disappears.
|
| 189 |
|
| 190 |
## Why it works
|
| 191 |
|
|
@@ -198,8 +194,9 @@ Most of lev's round trip from a laptop is network and Modal's ingress; its compu
|
|
| 198 |
### The optimizations that mattered
|
| 199 |
|
| 200 |
- **One batched forward instead of prefill-and-fork: 169 → 69 ms.** At this size the forward pass is bound by kernel launches, not arithmetic, so forking the cache saved FLOPs and cost time. One batched forward plus the depthwise-conv kernel cut compute by 59%. [ADR-023](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-023--one-batched-forward-not-prefill-and-fork)
|
| 201 |
-
- **Label-token readout up to the tokenizer's limit: +51 points on
|
| 202 |
- **Skipping codes that split: banking77 0.818 → 0.980.** Passing over codes that tokenize to two tokens lifts label-token readout from 68 options to several hundred. [ADR-028](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-028--skip-split-label-codes-when-serving-calibrate-for-families-the-model-has-not-seen)
|
|
|
|
| 203 |
|
| 204 |
## Training
|
| 205 |
|
|
@@ -209,18 +206,16 @@ Most of lev's round trip from a laptop is network and Modal's ingress; its compu
|
|
| 209 |
|
| 210 |
## Boundaries worth understanding
|
| 211 |
|
| 212 |
-
- **Minimal
|
| 213 |
-
- **Fine-grained quality ratings are weak.** Helpfulness scoring (helpsteer2) sits at 0.36. Treat such scores as a rough signal.
|
| 214 |
- **Calibration is fitted on the training distribution.** Temperatures are chosen to transfer across task families. Even so, a task very unlike the training mix may be less well calibrated. Check on your own data before you gate on the probabilities.
|
| 215 |
- **Questions are answered independently.** Answers in one request do not condition on each other. Encode a joint decision as one choice, or ask in stages.
|
| 216 |
-
- **Partial benchmark coverage.** S1Bench results cover 6 of its 13 subsets.
|
| 217 |
- **English only.**
|
| 218 |
- **Needs a GPU for real-time use.** It runs on CPU, but a 4B backbone there takes seconds per call, not milliseconds.
|
| 219 |
|
| 220 |
## Files
|
| 221 |
|
| 222 |
| file | role |
|
| 223 |
-
|---|---|
|
| 224 |
| `adapter_model.safetensors`, `adapter_config.json` | LoRA adapter |
|
| 225 |
| `mode_b_head.pt` | candidate-path head (tensor state dict, loaded with `weights_only=True`) |
|
| 226 |
| `tokenizer*`, `chat_template.jinja` | the tokenizer that the label codes were verified against |
|
|
|
|
| 26 |
|
| 27 |
<div align="center" style="line-height: 1;"><img src="https://img.shields.io/badge/license-Apache--2.0-2a78d6?style=flat-square" alt="Apache-2.0" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/base-Qwen3.5--4B-2a78d6?style=flat-square" alt="Qwen3.5-4B" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/output%20tokens-0-2a78d6?style=flat-square" alt="Zero output tokens" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/API-%2Fv1%2Fsystemone-2a78d6?style=flat-square" alt="/v1/systemone compatible" style="display: inline-block; vertical-align: middle; margin: 2px;"> <a href="https://github.com/Abhinavexists/lev"><img src="https://img.shields.io/badge/code-GitHub-14181f?style=flat-square&logo=github" alt="GitHub" style="display: inline-block; vertical-align: middle; margin: 2px;"></a></div>
|
| 28 |
|
| 29 |
+
<h2 align="center">68.9% on all 13 S1Bench subsets. 4B parameters. Zero output tokens.</h2>
|
| 30 |
|
| 31 |
+
<p align="center"><strong>Qwen3.5-4B + LoRA on one H100.</strong> On the six subsets the public S1Bench board completed, level with reflex-4b and behind only Jev and three open models of 26B–35B.</p>
|
| 32 |
|
| 33 |
+
<p align="center"><a href="#quickstart"><strong>Quickstart</strong></a> · <a href="#self-hosting-a-jev-compatible-http-server">Self-hosting</a> · <a href="#benchmarks">Benchmarks</a> · <a href="#speed">Speed</a> · <a href="#why-it-works">Why it works</a> · <a href="#training">Training</a> · <a href="#boundaries-worth-understanding">Boundaries</a> · <a href="https://github.com/Abhinavexists/lev">GitHub</a></p>
|
| 34 |
|
| 35 |
+
**No generated tokens, no JSON to parse, no retries.** lev cannot return a label outside your options, because the answer space is the option set you send. It can still pick the wrong option: the guarantee is structural, not a guarantee of correctness.
|
|
|
|
|
|
|
| 36 |
|
| 37 |
| Question | You give | You get |
|
| 38 |
+
| --- | --- | --- |
|
| 39 |
| `noul` | a yes/no question | `noul` = p(yes) |
|
| 40 |
| `choice` | instructions + options (name → description or `null`) | `choice`, `probabilities`, `confidence` |
|
| 41 |
| `score` | instructions + 2–10 ordered levels | `score` (expected level), `probabilities`, `confidence` |
|
|
|
|
| 131 |
|
| 132 |
### S1Bench
|
| 133 |
|
|
|
|
|
|
|
| 134 |
<p align="center"><img src="assets/accuracy-by-subset.png" alt="lev and Jev accuracy on each S1Bench subset" width="100%"></p>
|
| 135 |
|
| 136 |
+
| subset | task | **lev** | Jev | always the most common label |
|
| 137 |
+
| --- | --- | --: | --: | --: |
|
| 138 |
+
| vitaminc-dev | claim verification | 0.668 | **0.801** | 0.503 |
|
| 139 |
+
| massive-en-US | intent routing, 18 scenarios | 0.857 | **0.874** | 0.163 |
|
| 140 |
+
| massive-de-DE | intent routing, German | 0.823 | **0.871** | 0.163 |
|
| 141 |
+
| boolq | yes/no reading comprehension | 0.827 | **0.893** | 0.580 |
|
| 142 |
+
| squad2 | answerability | 0.813 | **0.836** | 0.502 |
|
| 143 |
+
| paws | adversarial paraphrase | 0.776 | **0.900** | 0.516 |
|
| 144 |
+
| multinli | natural language inference | **0.890** | 0.836 | 0.361 |
|
| 145 |
+
| civil_comments | toxicity | 0.760 | **0.803** | 0.893 |
|
| 146 |
+
| aegis2 | safety moderation | 0.800 | **0.804** | 0.568 |
|
| 147 |
+
| helpsteer2 | helpfulness, 5 levels | **0.386** | 0.341 | 0.422 |
|
| 148 |
+
| summeval-relevance | summary relevance, 5 levels | 0.358 | 0.358 | 0.458 |
|
| 149 |
+
| summeval-consistency | summary faithfulness, 5 levels | 0.271 | **0.812** | 0.840 |
|
| 150 |
+
| pubmedqa | biomedical yes/no/maybe | 0.732 | **0.764** | 0.532 |
|
| 151 |
+
| **macro** | | 0.689 | **0.761** | |
|
| 152 |
+
|
| 153 |
+
lev and TypeSafe Jev ran through the same harness on all 3,880 items S1Bench scores, pinned by [Nimble](https://github.com/bespokelabsai/nimble)'s manifests. Our Jev run lands within 0.8 points of TypeSafe's published figure on every subset, so the harness is not the gap.
|
| 154 |
+
|
| 155 |
+
- **Noise:** at these sizes a per-subset difference needs roughly 5–9 points to be real. lev's leads on multinli and helpsteer2 are inside that.
|
| 156 |
+
- **Where Jev is clearly ahead:** the minimal-edit pairs (paws −12.4, vitaminc −13.3) and summeval-consistency (−54.1), where lev rates most fully faithful summaries one level low. Fine-tuning introduced that: the untuned backbone scores 0.826.
|
| 157 |
+
- **Constant baselines:** on civil_comments (89% not toxic) and the three 5-level rating subsets, always answering the most common label beats both models.
|
| 158 |
+
- **Calibration:** mean ECE 0.115 for lev, 0.091 for Jev; lev is better calibrated on 5 of 13.
|
| 159 |
+
|
| 160 |
+
<p align="center"><img src="assets/leaderboard.png" alt="S1Bench leaderboard over the six subsets every listed model completed" width="100%"></p>
|
| 161 |
|
| 162 |
<p align="center"><img src="assets/accuracy-per-parameter.png" alt="Macro accuracy against parameter count" width="100%"></p>
|
| 163 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 164 |
### Held-out split
|
| 165 |
|
| 166 |
A held-out split of the 29 training sources, with no row shared with training:
|
| 167 |
|
| 168 |
| metric | value |
|
| 169 |
+
| --- | --- |
|
| 170 |
| weighted accuracy | 0.807\* |
|
| 171 |
| expected calibration error | 0.061\* (0.180 before calibration) |
|
| 172 |
| banking77 (77 intents) | 0.980 |
|
|
|
|
| 181 |
|
| 182 |
A call is one batched forward pass over every question, so compute stays flat from one question to eight, and a 60-option choice costs the same as a yes/no.
|
| 183 |
|
| 184 |
+
The 69 ms is engine compute for a short request (a three-sentence state), measured inside the container; S1Bench's longer states take more. End to end from a laptop, Jev's hosted API answered in 335–346 ms median and lev on one Modal H100 in 414–654 ms across two runs.
|
|
|
|
|
|
|
| 185 |
|
| 186 |
## Why it works
|
| 187 |
|
|
|
|
| 194 |
### The optimizations that mattered
|
| 195 |
|
| 196 |
- **One batched forward instead of prefill-and-fork: 169 → 69 ms.** At this size the forward pass is bound by kernel launches, not arithmetic, so forking the cache saved FLOPs and cost time. One batched forward plus the depthwise-conv kernel cut compute by 59%. [ADR-023](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-023--one-batched-forward-not-prefill-and-fork)
|
| 197 |
+
- **Label-token readout up to the tokenizer's limit: +51 points on MASSIVE's 60 intents.** Serving 60 options through label codes instead of the learned head took accuracy from 0.231 to 0.746. [ADR-025](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-025--serving-routes-mode-a-up-to-the-tokenizer-limit-training-keeps-its-cap)
|
| 198 |
- **Skipping codes that split: banking77 0.818 → 0.980.** Passing over codes that tokenize to two tokens lifts label-token readout from 68 options to several hundred. [ADR-028](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-028--skip-split-label-codes-when-serving-calibrate-for-families-the-model-has-not-seen)
|
| 199 |
+
- **The backbone's own prompt format: 0.653 → 0.710 frozen.** The untuned instruct model scored 5.7 points higher on the earlier six-subset S1Bench set when questions are dressed in its chat template, so the final run trained in that format. [ADR-027](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-027--the-prompt-is-dressed-in-the-backbones-own-format)
|
| 200 |
|
| 201 |
## Training
|
| 202 |
|
|
|
|
| 206 |
|
| 207 |
## Boundaries worth understanding
|
| 208 |
|
| 209 |
+
- **Minimal edits and fine-grained ratings are weak.** Inputs that differ by one swapped word or number, and quality ratings over five levels, are where lev is least accurate and can be confidently wrong.
|
|
|
|
| 210 |
- **Calibration is fitted on the training distribution.** Temperatures are chosen to transfer across task families. Even so, a task very unlike the training mix may be less well calibrated. Check on your own data before you gate on the probabilities.
|
| 211 |
- **Questions are answered independently.** Answers in one request do not condition on each other. Encode a joint decision as one choice, or ask in stages.
|
|
|
|
| 212 |
- **English only.**
|
| 213 |
- **Needs a GPU for real-time use.** It runs on CPU, but a 4B backbone there takes seconds per call, not milliseconds.
|
| 214 |
|
| 215 |
## Files
|
| 216 |
|
| 217 |
| file | role |
|
| 218 |
+
| --- | --- |
|
| 219 |
| `adapter_model.safetensors`, `adapter_config.json` | LoRA adapter |
|
| 220 |
| `mode_b_head.pt` | candidate-path head (tensor state dict, loaded with `weights_only=True`) |
|
| 221 |
| `tokenizer*`, `chat_template.jinja` | the tokenizer that the label codes were verified against |
|
assets/accuracy-by-subset.png
CHANGED
|
|
Git LFS Details
|
assets/accuracy-per-parameter.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
assets/compute.png
CHANGED
|
|
assets/latency.png
DELETED
|
Binary file (94.5 kB)
|
|
|
assets/leaderboard.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|