Abhinav-jigsawstack commited on
Commit
d045c05
·
verified ·
1 Parent(s): 4ac3955

docs: report lev vs Jev on all 13 S1Bench subsets

Browse files
.gitattributes CHANGED
@@ -37,3 +37,4 @@ assets/leaderboard.png filter=lfs diff=lfs merge=lfs -text
37
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
38
  assets/accuracy-per-parameter.png filter=lfs diff=lfs merge=lfs -text
39
  assets/hero.png filter=lfs diff=lfs merge=lfs -text
 
 
37
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
38
  assets/accuracy-per-parameter.png filter=lfs diff=lfs merge=lfs -text
39
  assets/hero.png filter=lfs diff=lfs merge=lfs -text
40
+ assets/accuracy-by-subset.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -26,18 +26,16 @@ lev answers typed questions about a piece of context in a single forward pass. Y
26
 
27
  <div align="center" style="line-height: 1;"><img src="https://img.shields.io/badge/license-Apache--2.0-2a78d6?style=flat-square" alt="Apache-2.0" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/base-Qwen3.5--4B-2a78d6?style=flat-square" alt="Qwen3.5-4B" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/output%20tokens-0-2a78d6?style=flat-square" alt="Zero output tokens" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/API-%2Fv1%2Fsystemone-2a78d6?style=flat-square" alt="/v1/systemone compatible" style="display: inline-block; vertical-align: middle; margin: 2px;"> <a href="https://github.com/Abhinavexists/lev"><img src="https://img.shields.io/badge/code-GitHub-14181f?style=flat-square&logo=github" alt="GitHub" style="display: inline-block; vertical-align: middle; margin: 2px;"></a></div>
28
 
29
- <h2 align="center">72.5% on S1Bench. 69 ms of compute. 4B parameters.</h2>
30
 
31
- <p align="center"><strong>Qwen3.5-4B + LoRA · one H100 · six S1Bench subsets, 1,999 items.</strong><br>Sixth of 22 on the S1Bench board over these subsets, ahead of every model its size or smaller.<br>69 ms is compute inside the container. End to end from a laptop it is 414–463 ms.<br><a href="#benchmarks">Benchmarks</a> · <a href="#speed">Speed</a> · <a href="#boundaries-worth-understanding">Boundaries</a></p>
32
 
33
- <p align="center"><a href="#quickstart"><strong>Quickstart</strong></a> · <a href="#self-hosting-a-jev-compatible-http-server">Self-hosting</a> · <a href="#benchmarks">Benchmarks</a> · <a href="#why-it-works">Why it works</a> · <a href="#training">Training</a> · <a href="https://github.com/Abhinavexists/lev">GitHub</a></p>
34
 
35
- **No generated tokens. No JSON to parse. No retries until it validates.**
36
-
37
- **lev cannot return a label outside your options.** The answer space is the option set you send, so every response is well-formed by construction. lev can still pick the wrong option: this is a structural guarantee, not a guarantee of correctness.
38
 
39
  | Question | You give | You get |
40
- |---|---|---|
41
  | `noul` | a yes/no question | `noul` = p(yes) |
42
  | `choice` | instructions + options (name → description or `null`) | `choice`, `probabilities`, `confidence` |
43
  | `score` | instructions + 2–10 ordered levels | `score` (expected level), `probabilities`, `confidence` |
@@ -133,42 +131,42 @@ print(response.answers["team"].choice, response.answers["bug"].noul) # technica
133
 
134
  ### S1Bench
135
 
136
- Six S1Bench subsets, 1,999 items. lev and TypeSafe Jev were run through the same harness on the same task files. Accuracy, best in each row in bold:
137
-
138
  <p align="center"><img src="assets/accuracy-by-subset.png" alt="lev and Jev accuracy on each S1Bench subset" width="100%"></p>
139
 
140
- | subset | task | **lev** | Jev | Qwen3.5-4B, untuned |
141
- |---|---|--:|--:|--:|
142
- | aegis2 | safety moderation | **0.864** | 0.832 | 0.776 |
143
- | boolq | yes/no reading comprehension | 0.880 | **0.910** | 0.860 |
144
- | massive-en-US | intent, 60 classes | 0.791 | **0.814** | 0.734 |
145
- | vitaminc-dev | claim verification | 0.738 | **0.846** | 0.733 |
146
- | paws | adversarial paraphrase | 0.716 | **0.820** | 0.756 |
147
- | helpsteer2 | helpfulness rating | 0.360 | 0.304 | **0.400** |
148
- | **macro** | | 0.725 | **0.754** | 0.710 |
149
-
150
- > **What the headline means:** accuracy on 6 of S1Bench's 13 subsets, not all of them. At these sizes each subset figure carries roughly ±3–5 points. Our harness scores Jev 2.1 points below its figure on the public board, so lev's placement there is conservative. Jev leads on macro accuracy.
151
-
152
- <p align="center"><img src="assets/leaderboard.png" alt="S1Bench leaderboard: lev ranks sixth of 22" width="100%"></p>
153
-
154
- On the S1Bench board over these subsets, lev ranks sixth of 22. Only Jev and three open models of 26B–35B parameters score higher.
 
 
 
 
 
 
 
 
 
 
155
 
156
  <p align="center"><img src="assets/accuracy-per-parameter.png" alt="Macro accuracy against parameter count" width="100%"></p>
157
 
158
- #### Where Jev leads
159
-
160
- - **Minimal-edit pairs:** paws (−10.4 points) and vitaminc (−10.8), where two inputs differ by one swapped word or one changed number.
161
- - **boolq and 60-class intent:** 3.0 and 2.3 points behind.
162
- - **End-to-end latency from a laptop:** Jev's hosted API answered in 344–357 ms median, lev on one Modal H100 in 414–463 ms.
163
-
164
- lev leads on safety moderation (+3.2) and helpfulness rating (+5.6).
165
-
166
  ### Held-out split
167
 
168
  A held-out split of the 29 training sources, with no row shared with training:
169
 
170
  | metric | value |
171
- |---|---|
172
  | weighted accuracy | 0.807\* |
173
  | expected calibration error | 0.061\* (0.180 before calibration) |
174
  | banking77 (77 intents) | 0.980 |
@@ -183,9 +181,7 @@ A held-out split of the 29 training sources, with no row shared with training:
183
 
184
  A call is one batched forward pass over every question, so compute stays flat from one question to eight, and a 60-option choice costs the same as a yes/no.
185
 
186
- <p align="center"><img src="assets/latency.png" alt="Per-call latency from the same laptop, lev and Jev" width="100%"></p>
187
-
188
- Most of lev's round trip from a laptop is network and Modal's ingress; its compute is 69 ms. Self-hosted next to your application, that network hop disappears.
189
 
190
  ## Why it works
191
 
@@ -198,8 +194,9 @@ Most of lev's round trip from a laptop is network and Modal's ingress; its compu
198
  ### The optimizations that mattered
199
 
200
  - **One batched forward instead of prefill-and-fork: 169 → 69 ms.** At this size the forward pass is bound by kernel launches, not arithmetic, so forking the cache saved FLOPs and cost time. One batched forward plus the depthwise-conv kernel cut compute by 59%. [ADR-023](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-023--one-batched-forward-not-prefill-and-fork)
201
- - **Label-token readout up to the tokenizer's limit: +51 points on massive-en-US.** Serving 60 options through label codes instead of the learned head took 60-class intent from 0.231 to 0.746. [ADR-025](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-025--serving-routes-mode-a-up-to-the-tokenizer-limit-training-keeps-its-cap)
202
  - **Skipping codes that split: banking77 0.818 → 0.980.** Passing over codes that tokenize to two tokens lifts label-token readout from 68 options to several hundred. [ADR-028](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-028--skip-split-label-codes-when-serving-calibrate-for-families-the-model-has-not-seen)
 
203
 
204
  ## Training
205
 
@@ -209,18 +206,16 @@ Most of lev's round trip from a laptop is network and Modal's ingress; its compu
209
 
210
  ## Boundaries worth understanding
211
 
212
- - **Minimal-edit pairs.** On paws and vitaminc, two inputs can differ by one swapped word or one changed number. Here the model can be confidently wrong. It scores below the untuned backbone on paws and only matches it on vitaminc.
213
- - **Fine-grained quality ratings are weak.** Helpfulness scoring (helpsteer2) sits at 0.36. Treat such scores as a rough signal.
214
  - **Calibration is fitted on the training distribution.** Temperatures are chosen to transfer across task families. Even so, a task very unlike the training mix may be less well calibrated. Check on your own data before you gate on the probabilities.
215
  - **Questions are answered independently.** Answers in one request do not condition on each other. Encode a joint decision as one choice, or ask in stages.
216
- - **Partial benchmark coverage.** S1Bench results cover 6 of its 13 subsets.
217
  - **English only.**
218
  - **Needs a GPU for real-time use.** It runs on CPU, but a 4B backbone there takes seconds per call, not milliseconds.
219
 
220
  ## Files
221
 
222
  | file | role |
223
- |---|---|
224
  | `adapter_model.safetensors`, `adapter_config.json` | LoRA adapter |
225
  | `mode_b_head.pt` | candidate-path head (tensor state dict, loaded with `weights_only=True`) |
226
  | `tokenizer*`, `chat_template.jinja` | the tokenizer that the label codes were verified against |
 
26
 
27
  <div align="center" style="line-height: 1;"><img src="https://img.shields.io/badge/license-Apache--2.0-2a78d6?style=flat-square" alt="Apache-2.0" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/base-Qwen3.5--4B-2a78d6?style=flat-square" alt="Qwen3.5-4B" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/output%20tokens-0-2a78d6?style=flat-square" alt="Zero output tokens" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img src="https://img.shields.io/badge/API-%2Fv1%2Fsystemone-2a78d6?style=flat-square" alt="/v1/systemone compatible" style="display: inline-block; vertical-align: middle; margin: 2px;"> <a href="https://github.com/Abhinavexists/lev"><img src="https://img.shields.io/badge/code-GitHub-14181f?style=flat-square&logo=github" alt="GitHub" style="display: inline-block; vertical-align: middle; margin: 2px;"></a></div>
28
 
29
+ <h2 align="center">68.9% on all 13 S1Bench subsets. 4B parameters. Zero output tokens.</h2>
30
 
31
+ <p align="center"><strong>Qwen3.5-4B + LoRA on one H100.</strong> On the six subsets the public S1Bench board completed, level with reflex-4b and behind only Jev and three open models of 26B–35B.</p>
32
 
33
+ <p align="center"><a href="#quickstart"><strong>Quickstart</strong></a> · <a href="#self-hosting-a-jev-compatible-http-server">Self-hosting</a> · <a href="#benchmarks">Benchmarks</a> · <a href="#speed">Speed</a> · <a href="#why-it-works">Why it works</a> · <a href="#training">Training</a> · <a href="#boundaries-worth-understanding">Boundaries</a> · <a href="https://github.com/Abhinavexists/lev">GitHub</a></p>
34
 
35
+ **No generated tokens, no JSON to parse, no retries.** lev cannot return a label outside your options, because the answer space is the option set you send. It can still pick the wrong option: the guarantee is structural, not a guarantee of correctness.
 
 
36
 
37
  | Question | You give | You get |
38
+ | --- | --- | --- |
39
  | `noul` | a yes/no question | `noul` = p(yes) |
40
  | `choice` | instructions + options (name → description or `null`) | `choice`, `probabilities`, `confidence` |
41
  | `score` | instructions + 2–10 ordered levels | `score` (expected level), `probabilities`, `confidence` |
 
131
 
132
  ### S1Bench
133
 
 
 
134
  <p align="center"><img src="assets/accuracy-by-subset.png" alt="lev and Jev accuracy on each S1Bench subset" width="100%"></p>
135
 
136
+ | subset | task | **lev** | Jev | always the most common label |
137
+ | --- | --- | --: | --: | --: |
138
+ | vitaminc-dev | claim verification | 0.668 | **0.801** | 0.503 |
139
+ | massive-en-US | intent routing, 18 scenarios | 0.857 | **0.874** | 0.163 |
140
+ | massive-de-DE | intent routing, German | 0.823 | **0.871** | 0.163 |
141
+ | boolq | yes/no reading comprehension | 0.827 | **0.893** | 0.580 |
142
+ | squad2 | answerability | 0.813 | **0.836** | 0.502 |
143
+ | paws | adversarial paraphrase | 0.776 | **0.900** | 0.516 |
144
+ | multinli | natural language inference | **0.890** | 0.836 | 0.361 |
145
+ | civil_comments | toxicity | 0.760 | **0.803** | 0.893 |
146
+ | aegis2 | safety moderation | 0.800 | **0.804** | 0.568 |
147
+ | helpsteer2 | helpfulness, 5 levels | **0.386** | 0.341 | 0.422 |
148
+ | summeval-relevance | summary relevance, 5 levels | 0.358 | 0.358 | 0.458 |
149
+ | summeval-consistency | summary faithfulness, 5 levels | 0.271 | **0.812** | 0.840 |
150
+ | pubmedqa | biomedical yes/no/maybe | 0.732 | **0.764** | 0.532 |
151
+ | **macro** | | 0.689 | **0.761** | |
152
+
153
+ lev and TypeSafe Jev ran through the same harness on all 3,880 items S1Bench scores, pinned by [Nimble](https://github.com/bespokelabsai/nimble)'s manifests. Our Jev run lands within 0.8 points of TypeSafe's published figure on every subset, so the harness is not the gap.
154
+
155
+ - **Noise:** at these sizes a per-subset difference needs roughly 5–9 points to be real. lev's leads on multinli and helpsteer2 are inside that.
156
+ - **Where Jev is clearly ahead:** the minimal-edit pairs (paws −12.4, vitaminc −13.3) and summeval-consistency (−54.1), where lev rates most fully faithful summaries one level low. Fine-tuning introduced that: the untuned backbone scores 0.826.
157
+ - **Constant baselines:** on civil_comments (89% not toxic) and the three 5-level rating subsets, always answering the most common label beats both models.
158
+ - **Calibration:** mean ECE 0.115 for lev, 0.091 for Jev; lev is better calibrated on 5 of 13.
159
+
160
+ <p align="center"><img src="assets/leaderboard.png" alt="S1Bench leaderboard over the six subsets every listed model completed" width="100%"></p>
161
 
162
  <p align="center"><img src="assets/accuracy-per-parameter.png" alt="Macro accuracy against parameter count" width="100%"></p>
163
 
 
 
 
 
 
 
 
 
164
  ### Held-out split
165
 
166
  A held-out split of the 29 training sources, with no row shared with training:
167
 
168
  | metric | value |
169
+ | --- | --- |
170
  | weighted accuracy | 0.807\* |
171
  | expected calibration error | 0.061\* (0.180 before calibration) |
172
  | banking77 (77 intents) | 0.980 |
 
181
 
182
  A call is one batched forward pass over every question, so compute stays flat from one question to eight, and a 60-option choice costs the same as a yes/no.
183
 
184
+ The 69 ms is engine compute for a short request (a three-sentence state), measured inside the container; S1Bench's longer states take more. End to end from a laptop, Jev's hosted API answered in 335–346 ms median and lev on one Modal H100 in 414–654 ms across two runs.
 
 
185
 
186
  ## Why it works
187
 
 
194
  ### The optimizations that mattered
195
 
196
  - **One batched forward instead of prefill-and-fork: 169 → 69 ms.** At this size the forward pass is bound by kernel launches, not arithmetic, so forking the cache saved FLOPs and cost time. One batched forward plus the depthwise-conv kernel cut compute by 59%. [ADR-023](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-023--one-batched-forward-not-prefill-and-fork)
197
+ - **Label-token readout up to the tokenizer's limit: +51 points on MASSIVE's 60 intents.** Serving 60 options through label codes instead of the learned head took accuracy from 0.231 to 0.746. [ADR-025](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-025--serving-routes-mode-a-up-to-the-tokenizer-limit-training-keeps-its-cap)
198
  - **Skipping codes that split: banking77 0.818 → 0.980.** Passing over codes that tokenize to two tokens lifts label-token readout from 68 options to several hundred. [ADR-028](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-028--skip-split-label-codes-when-serving-calibrate-for-families-the-model-has-not-seen)
199
+ - **The backbone's own prompt format: 0.653 → 0.710 frozen.** The untuned instruct model scored 5.7 points higher on the earlier six-subset S1Bench set when questions are dressed in its chat template, so the final run trained in that format. [ADR-027](https://github.com/Abhinavexists/lev/blob/main/docs/DECISIONS.md#adr-027--the-prompt-is-dressed-in-the-backbones-own-format)
200
 
201
  ## Training
202
 
 
206
 
207
  ## Boundaries worth understanding
208
 
209
+ - **Minimal edits and fine-grained ratings are weak.** Inputs that differ by one swapped word or number, and quality ratings over five levels, are where lev is least accurate and can be confidently wrong.
 
210
  - **Calibration is fitted on the training distribution.** Temperatures are chosen to transfer across task families. Even so, a task very unlike the training mix may be less well calibrated. Check on your own data before you gate on the probabilities.
211
  - **Questions are answered independently.** Answers in one request do not condition on each other. Encode a joint decision as one choice, or ask in stages.
 
212
  - **English only.**
213
  - **Needs a GPU for real-time use.** It runs on CPU, but a 4B backbone there takes seconds per call, not milliseconds.
214
 
215
  ## Files
216
 
217
  | file | role |
218
+ | --- | --- |
219
  | `adapter_model.safetensors`, `adapter_config.json` | LoRA adapter |
220
  | `mode_b_head.pt` | candidate-path head (tensor state dict, loaded with `weights_only=True`) |
221
  | `tokenizer*`, `chat_template.jinja` | the tokenizer that the label codes were verified against |
assets/accuracy-by-subset.png CHANGED

Git LFS Details

  • SHA256: 2d5941819910e61034af2a48ac5800db2409eaa2ef786452b577477179c3e1b3
  • Pointer size: 131 Bytes
  • Size of remote file: 176 kB
assets/accuracy-per-parameter.png CHANGED

Git LFS Details

  • SHA256: dac47e9c68b2aadeaff10217b037c226f5b0d4a0fed7311bb97bb7dcc7a2dab6
  • Pointer size: 131 Bytes
  • Size of remote file: 103 kB

Git LFS Details

  • SHA256: ed5465d6df019c81594c05434b7b1e890f5f8e696562dbc97d2b91f92f504022
  • Pointer size: 131 Bytes
  • Size of remote file: 103 kB
assets/compute.png CHANGED
assets/latency.png DELETED
Binary file (94.5 kB)
 
assets/leaderboard.png CHANGED

Git LFS Details

  • SHA256: 23735f3ea5bd43fd39e22996b3ff14919209c669760eed930239a4ba7910a10e
  • Pointer size: 131 Bytes
  • Size of remote file: 241 kB

Git LFS Details

  • SHA256: 0f758b2d7cc5638cff05dd34d498dc82fa0e4f19ab6ad7ecc5d22f0188a9dc57
  • Pointer size: 131 Bytes
  • Size of remote file: 241 kB