Instructions to use HopitAI/hopper with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use HopitAI/hopper with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B") model = PeftModel.from_pretrained(base_model, "HopitAI/hopper") - Notebooks
- Google Colab
- Kaggle
Hopper v1.1.0: calibration map replaced by one temperature per answer type; adapter weights unchanged
Browse files- README.md +31 -10
- hopper.json +12 -38
README.md
CHANGED
|
@@ -37,10 +37,16 @@ question go in, and a probability distribution over a fixed set of options comes
|
|
| 37 |
- **One forward pass per decision**, with thinking off. No text is generated. The answer is a
|
| 38 |
softmax over the logits of the option letters (A, B, C, ...), restricted to as many letters as
|
| 39 |
there are options.
|
| 40 |
-
- **A calibration map** (`hopper.json`) rescales that distribution by
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 44 |
- **Serving code**: [github.com/hopit-ai/hopper](https://github.com/hopit-ai/hopper). It runs an
|
| 45 |
HTTP server with the JevBench `/v1/systemone` wire format. At load it merges the adapter into the
|
| 46 |
bf16 weights, and it refuses to start if the fast linear-attention kernels are not active.
|
|
@@ -134,8 +140,8 @@ The adapter was trained on a mix of three sources:
|
|
| 134 |
| [fancyzhx/dbpedia_14](https://huggingface.co/datasets/fancyzhx/dbpedia_14) | topic classification | CC BY-SA 3.0 |
|
| 135 |
| [nvidia/HelpSteer2](https://huggingface.co/datasets/nvidia/HelpSteer2) | response-quality judgement | CC BY 4.0 |
|
| 136 |
|
| 137 |
-
The calibration map was fitted only on our own held-out JevBench-style items
|
| 138 |
-
JevBench item.
|
| 139 |
|
| 140 |
## Evaluation
|
| 141 |
|
|
@@ -151,8 +157,19 @@ official.
|
|
| 151 |
| standard (original) | 72 | 0.944 | 0.958 |
|
| 152 |
| hard | 111 | 0.685 | 0.631 |
|
| 153 |
|
| 154 |
-
|
| 155 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 156 |
|
| 157 |
**Disclosure.**
|
| 158 |
- The public items were split in half before we started. The half we developed on (115 items)
|
|
@@ -165,7 +182,9 @@ distance) on the 10 public probability items is 0.830.
|
|
| 165 |
chosen beforehand by a pre-registered rule. On it the adapter scores hard 0.661 (37 of 56) and
|
| 166 |
the frozen base 0.643 (36 of 56): it is level with the frozen model on accuracy there, not ahead.
|
| 167 |
- The calibration map was fitted only on our own held-out JevBench-style items, never on a
|
| 168 |
-
JevBench item.
|
|
|
|
|
|
|
| 169 |
- No JevBench item or paraphrase was used in training. Every item we wrote was checked against
|
| 170 |
all public JevBench questions and states (normalised question identity, and any shared 8-word
|
| 171 |
sequence) and dropped on a match; the check reads only hashes and reports only counts.
|
|
@@ -181,7 +200,9 @@ distance) on the 10 public probability items is 0.830.
|
|
| 181 |
- **Long documents that need several hops are weak.** Accuracy drops when the answer needs facts
|
| 182 |
from several distant parts of a long document.
|
| 183 |
- **The calibration was fitted on our own data.** The map was fitted on our own JevBench-style
|
| 184 |
-
items. On a different distribution of questions, its confidences can be off.
|
|
|
|
|
|
|
| 185 |
- **The dev half flatters it.** On the reserved half of the public items, it is level with the
|
| 186 |
frozen base model on hard-tier accuracy (see the disclosure).
|
| 187 |
- **Tested only on English.** We have not measured any other language.
|
|
|
|
| 37 |
- **One forward pass per decision**, with thinking off. No text is generated. The answer is a
|
| 38 |
softmax over the logits of the option letters (A, B, C, ...), restricted to as many letters as
|
| 39 |
there are options.
|
| 40 |
+
- **A calibration map** (`hopper.json`) rescales that distribution by one temperature per answer
|
| 41 |
+
type: choice 0.790, noul 0.753, score 0.900. It was fitted only on our own held-out
|
| 42 |
+
JevBench-style items, never on a JevBench item. The map never changes the top answer — it
|
| 43 |
+
divides log-probabilities by a positive number, which cannot reorder them.
|
| 44 |
+
(In 1.0.0 the map was instead a bounded linear function of option count, state length,
|
| 45 |
+
JSON-or-not, answer type and the entropy of the model's own distribution. A leave-one-source-out
|
| 46 |
+
ablation showed that map was worth +0.1 Calibration over no map at all on sources its fitting
|
| 47 |
+
set had never seen, against +4.5 and +2.9 for the per-answer-type map fitted on the same data,
|
| 48 |
+
so 1.1.0 replaced it. Answers are identical either way; only the confidences move. The old map
|
| 49 |
+
ships alongside as `hopper-v1.0-linear.json`.)
|
| 50 |
- **Serving code**: [github.com/hopit-ai/hopper](https://github.com/hopit-ai/hopper). It runs an
|
| 51 |
HTTP server with the JevBench `/v1/systemone` wire format. At load it merges the adapter into the
|
| 52 |
bf16 weights, and it refuses to start if the fast linear-attention kernels are not active.
|
|
|
|
| 140 |
| [fancyzhx/dbpedia_14](https://huggingface.co/datasets/fancyzhx/dbpedia_14) | topic classification | CC BY-SA 3.0 |
|
| 141 |
| [nvidia/HelpSteer2](https://huggingface.co/datasets/nvidia/HelpSteer2) | response-quality judgement | CC BY 4.0 |
|
| 142 |
|
| 143 |
+
The calibration map was fitted only on our own held-out JevBench-style items — the same fitting
|
| 144 |
+
set in 1.1.0 as in 1.0.0, with the 66 leakage-flagged items dropped. It never saw a JevBench item.
|
| 145 |
|
| 146 |
## Evaluation
|
| 147 |
|
|
|
|
| 157 |
| standard (original) | 72 | 0.944 | 0.958 |
|
| 158 |
| hard | 111 | 0.685 | 0.631 |
|
| 159 |
|
| 160 |
+
The accuracies above are the same in 1.1.0 as in 1.0.0: no calibration map can move an answer.
|
| 161 |
+
|
| 162 |
+
On the hard tier, top-label ECE is 0.102, and distribution fidelity (1 − mean total-variation
|
| 163 |
+
distance) on the 10 public probability items is 0.830. **Both were measured with the 1.0.0
|
| 164 |
+
calibration map**, and we have not recomputed them for 1.1.0, because the reserved half of the
|
| 165 |
+
public items may be scored only under the pre-registered rule and we will not spend it again on
|
| 166 |
+
a map change. On the development half alone, computed on our saved predictions, replacing the
|
| 167 |
+
1.0.0 map with the 1.1.0 one takes hard-tier ECE from 0.159 to 0.112 and the harness's Calibration
|
| 168 |
+
axis from 76.5 to 79.4, with the top answers unchanged. Treat that as an estimate and not a
|
| 169 |
+
promise: re-running the same 55 hard items through the server applies the identical map and lands
|
| 170 |
+
at 76.0, because a binned ECE over 55 items is not stable to the bf16 differences between two runs
|
| 171 |
+
of the same weights. The case for the map rests on a 517-item held-out set, where it is worth +4.5,
|
| 172 |
+
not on this fold.
|
| 173 |
|
| 174 |
**Disclosure.**
|
| 175 |
- The public items were split in half before we started. The half we developed on (115 items)
|
|
|
|
| 182 |
chosen beforehand by a pre-registered rule. On it the adapter scores hard 0.661 (37 of 56) and
|
| 183 |
the frozen base 0.643 (36 of 56): it is level with the frozen model on accuracy there, not ahead.
|
| 184 |
- The calibration map was fitted only on our own held-out JevBench-style items, never on a
|
| 185 |
+
JevBench item. The 1.1.0 map was fitted on exactly the same items as the 1.0.0 map; it was
|
| 186 |
+
chosen over it on out-of-pool folds of our own data, not on any JevBench score, and the
|
| 187 |
+
reserved half was not re-run for it.
|
| 188 |
- No JevBench item or paraphrase was used in training. Every item we wrote was checked against
|
| 189 |
all public JevBench questions and states (normalised question identity, and any shared 8-word
|
| 190 |
sequence) and dropped on a match; the check reads only hashes and reports only counts.
|
|
|
|
| 200 |
- **Long documents that need several hops are weak.** Accuracy drops when the answer needs facts
|
| 201 |
from several distant parts of a long document.
|
| 202 |
- **The calibration was fitted on our own data.** The map was fitted on our own JevBench-style
|
| 203 |
+
items. On a different distribution of questions, its confidences can be off. 1.1.0 made the map
|
| 204 |
+
as simple as we could justify — three temperatures instead of seven coefficients — precisely
|
| 205 |
+
because the richer map did not transfer off the pool it was fitted on.
|
| 206 |
- **The dev half flatters it.** On the reserved half of the public items, it is level with the
|
| 207 |
frozen base model on hard-tier accuracy (see the disclosure).
|
| 208 |
- **Tested only on English.** We have not measured any other language.
|
hopper.json
CHANGED
|
@@ -1,40 +1,14 @@
|
|
| 1 |
{
|
| 2 |
-
"kind": "
|
| 3 |
-
"
|
| 4 |
-
|
| 5 |
-
"
|
| 6 |
-
"
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
"
|
| 10 |
-
"
|
| 11 |
-
"
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
-0.19086088298164494,
|
| 15 |
-
-0.5838464772898263,
|
| 16 |
-
-0.014374404205570545,
|
| 17 |
-
-0.036048925682095466,
|
| 18 |
-
-0.5757853955715236,
|
| 19 |
-
0.0032798295050555613,
|
| 20 |
-
0.19385182563994008
|
| 21 |
-
],
|
| 22 |
-
"means": [
|
| 23 |
-
0.0,
|
| 24 |
-
1.194336919379479,
|
| 25 |
-
5.574204769595688,
|
| 26 |
-
0.3476482617586912,
|
| 27 |
-
0.294478527607362,
|
| 28 |
-
0.05521472392638037,
|
| 29 |
-
0.6346005868967624
|
| 30 |
-
],
|
| 31 |
-
"deviations": [
|
| 32 |
-
1.0,
|
| 33 |
-
0.354090510000433,
|
| 34 |
-
0.7115462952612505,
|
| 35 |
-
0.47622363218854624,
|
| 36 |
-
0.45580799069955114,
|
| 37 |
-
0.2283989014599544,
|
| 38 |
-
0.19893323309471178
|
| 39 |
-
]
|
| 40 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"kind": "per_kind",
|
| 3 |
+
"temperatures": {
|
| 4 |
+
"choice": 0.7898505975532796,
|
| 5 |
+
"noul": 0.7531284680253558,
|
| 6 |
+
"score": 0.8997033695617845
|
| 7 |
+
},
|
| 8 |
+
"fitted_on": {
|
| 9 |
+
"description": "our own held-out JevBench-style look-alike items, hard tier: three measurement files, 622 items, 536 on the axis. No JevBench item and no JevBench file.",
|
| 10 |
+
"tier": "hard",
|
| 11 |
+
"map": "per_kind",
|
| 12 |
+
"dropped_leakage_flagged": 66
|
| 13 |
+
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 14 |
}
|