ArkidMitra commited on
Commit
281d393
·
verified ·
1 Parent(s): be8d454

Hopper v1.1.0: calibration map replaced by one temperature per answer type; adapter weights unchanged

Browse files
Files changed (2) hide show
  1. README.md +31 -10
  2. hopper.json +12 -38
README.md CHANGED
@@ -37,10 +37,16 @@ question go in, and a probability distribution over a fixed set of options comes
37
  - **One forward pass per decision**, with thinking off. No text is generated. The answer is a
38
  softmax over the logits of the option letters (A, B, C, ...), restricted to as many letters as
39
  there are options.
40
- - **A calibration map** (`hopper.json`) rescales that distribution by a temperature, T in
41
- [1/3, 3]. T is a bounded linear function of what the request shows: the number of options, the
42
- state length, whether the state is JSON, the answer type, and the entropy of the model's own
43
- distribution. The map never changes the top answer.
 
 
 
 
 
 
44
  - **Serving code**: [github.com/hopit-ai/hopper](https://github.com/hopit-ai/hopper). It runs an
45
  HTTP server with the JevBench `/v1/systemone` wire format. At load it merges the adapter into the
46
  bf16 weights, and it refuses to start if the fast linear-attention kernels are not active.
@@ -134,8 +140,8 @@ The adapter was trained on a mix of three sources:
134
  | [fancyzhx/dbpedia_14](https://huggingface.co/datasets/fancyzhx/dbpedia_14) | topic classification | CC BY-SA 3.0 |
135
  | [nvidia/HelpSteer2](https://huggingface.co/datasets/nvidia/HelpSteer2) | response-quality judgement | CC BY 4.0 |
136
 
137
- The calibration map was fitted only on our own held-out JevBench-style items. It never saw a
138
- JevBench item.
139
 
140
  ## Evaluation
141
 
@@ -151,8 +157,19 @@ official.
151
  | standard (original) | 72 | 0.944 | 0.958 |
152
  | hard | 111 | 0.685 | 0.631 |
153
 
154
- On the hard tier, top-label ECE is 0.102. Distribution fidelity (1 − mean total-variation
155
- distance) on the 10 public probability items is 0.830.
 
 
 
 
 
 
 
 
 
 
 
156
 
157
  **Disclosure.**
158
  - The public items were split in half before we started. The half we developed on (115 items)
@@ -165,7 +182,9 @@ distance) on the 10 public probability items is 0.830.
165
  chosen beforehand by a pre-registered rule. On it the adapter scores hard 0.661 (37 of 56) and
166
  the frozen base 0.643 (36 of 56): it is level with the frozen model on accuracy there, not ahead.
167
  - The calibration map was fitted only on our own held-out JevBench-style items, never on a
168
- JevBench item.
 
 
169
  - No JevBench item or paraphrase was used in training. Every item we wrote was checked against
170
  all public JevBench questions and states (normalised question identity, and any shared 8-word
171
  sequence) and dropped on a match; the check reads only hashes and reports only counts.
@@ -181,7 +200,9 @@ distance) on the 10 public probability items is 0.830.
181
  - **Long documents that need several hops are weak.** Accuracy drops when the answer needs facts
182
  from several distant parts of a long document.
183
  - **The calibration was fitted on our own data.** The map was fitted on our own JevBench-style
184
- items. On a different distribution of questions, its confidences can be off.
 
 
185
  - **The dev half flatters it.** On the reserved half of the public items, it is level with the
186
  frozen base model on hard-tier accuracy (see the disclosure).
187
  - **Tested only on English.** We have not measured any other language.
 
37
  - **One forward pass per decision**, with thinking off. No text is generated. The answer is a
38
  softmax over the logits of the option letters (A, B, C, ...), restricted to as many letters as
39
  there are options.
40
+ - **A calibration map** (`hopper.json`) rescales that distribution by one temperature per answer
41
+ type: choice 0.790, noul 0.753, score 0.900. It was fitted only on our own held-out
42
+ JevBench-style items, never on a JevBench item. The map never changes the top answer — it
43
+ divides log-probabilities by a positive number, which cannot reorder them.
44
+ (In 1.0.0 the map was instead a bounded linear function of option count, state length,
45
+ JSON-or-not, answer type and the entropy of the model's own distribution. A leave-one-source-out
46
+ ablation showed that map was worth +0.1 Calibration over no map at all on sources its fitting
47
+ set had never seen, against +4.5 and +2.9 for the per-answer-type map fitted on the same data,
48
+ so 1.1.0 replaced it. Answers are identical either way; only the confidences move. The old map
49
+ ships alongside as `hopper-v1.0-linear.json`.)
50
  - **Serving code**: [github.com/hopit-ai/hopper](https://github.com/hopit-ai/hopper). It runs an
51
  HTTP server with the JevBench `/v1/systemone` wire format. At load it merges the adapter into the
52
  bf16 weights, and it refuses to start if the fast linear-attention kernels are not active.
 
140
  | [fancyzhx/dbpedia_14](https://huggingface.co/datasets/fancyzhx/dbpedia_14) | topic classification | CC BY-SA 3.0 |
141
  | [nvidia/HelpSteer2](https://huggingface.co/datasets/nvidia/HelpSteer2) | response-quality judgement | CC BY 4.0 |
142
 
143
+ The calibration map was fitted only on our own held-out JevBench-style items — the same fitting
144
+ set in 1.1.0 as in 1.0.0, with the 66 leakage-flagged items dropped. It never saw a JevBench item.
145
 
146
  ## Evaluation
147
 
 
157
  | standard (original) | 72 | 0.944 | 0.958 |
158
  | hard | 111 | 0.685 | 0.631 |
159
 
160
+ The accuracies above are the same in 1.1.0 as in 1.0.0: no calibration map can move an answer.
161
+
162
+ On the hard tier, top-label ECE is 0.102, and distribution fidelity (1 − mean total-variation
163
+ distance) on the 10 public probability items is 0.830. **Both were measured with the 1.0.0
164
+ calibration map**, and we have not recomputed them for 1.1.0, because the reserved half of the
165
+ public items may be scored only under the pre-registered rule and we will not spend it again on
166
+ a map change. On the development half alone, computed on our saved predictions, replacing the
167
+ 1.0.0 map with the 1.1.0 one takes hard-tier ECE from 0.159 to 0.112 and the harness's Calibration
168
+ axis from 76.5 to 79.4, with the top answers unchanged. Treat that as an estimate and not a
169
+ promise: re-running the same 55 hard items through the server applies the identical map and lands
170
+ at 76.0, because a binned ECE over 55 items is not stable to the bf16 differences between two runs
171
+ of the same weights. The case for the map rests on a 517-item held-out set, where it is worth +4.5,
172
+ not on this fold.
173
 
174
  **Disclosure.**
175
  - The public items were split in half before we started. The half we developed on (115 items)
 
182
  chosen beforehand by a pre-registered rule. On it the adapter scores hard 0.661 (37 of 56) and
183
  the frozen base 0.643 (36 of 56): it is level with the frozen model on accuracy there, not ahead.
184
  - The calibration map was fitted only on our own held-out JevBench-style items, never on a
185
+ JevBench item. The 1.1.0 map was fitted on exactly the same items as the 1.0.0 map; it was
186
+ chosen over it on out-of-pool folds of our own data, not on any JevBench score, and the
187
+ reserved half was not re-run for it.
188
  - No JevBench item or paraphrase was used in training. Every item we wrote was checked against
189
  all public JevBench questions and states (normalised question identity, and any shared 8-word
190
  sequence) and dropped on a match; the check reads only hashes and reports only counts.
 
200
  - **Long documents that need several hops are weak.** Accuracy drops when the answer needs facts
201
  from several distant parts of a long document.
202
  - **The calibration was fitted on our own data.** The map was fitted on our own JevBench-style
203
+ items. On a different distribution of questions, its confidences can be off. 1.1.0 made the map
204
+ as simple as we could justify — three temperatures instead of seven coefficients — precisely
205
+ because the richer map did not transfer off the pool it was fitted on.
206
  - **The dev half flatters it.** On the reserved half of the public items, it is level with the
207
  frozen base model on hard-tier accuracy (see the disclosure).
208
  - **Tested only on English.** We have not measured any other language.
hopper.json CHANGED
@@ -1,40 +1,14 @@
1
  {
2
- "kind": "linear",
3
- "bound": 1.0986122886681098,
4
- "names": [
5
- "intercept",
6
- "log_options",
7
- "log_words",
8
- "json_state",
9
- "is_noul",
10
- "is_score",
11
- "entropy"
12
- ],
13
- "weights": [
14
- -0.19086088298164494,
15
- -0.5838464772898263,
16
- -0.014374404205570545,
17
- -0.036048925682095466,
18
- -0.5757853955715236,
19
- 0.0032798295050555613,
20
- 0.19385182563994008
21
- ],
22
- "means": [
23
- 0.0,
24
- 1.194336919379479,
25
- 5.574204769595688,
26
- 0.3476482617586912,
27
- 0.294478527607362,
28
- 0.05521472392638037,
29
- 0.6346005868967624
30
- ],
31
- "deviations": [
32
- 1.0,
33
- 0.354090510000433,
34
- 0.7115462952612505,
35
- 0.47622363218854624,
36
- 0.45580799069955114,
37
- 0.2283989014599544,
38
- 0.19893323309471178
39
- ]
40
  }
 
1
  {
2
+ "kind": "per_kind",
3
+ "temperatures": {
4
+ "choice": 0.7898505975532796,
5
+ "noul": 0.7531284680253558,
6
+ "score": 0.8997033695617845
7
+ },
8
+ "fitted_on": {
9
+ "description": "our own held-out JevBench-style look-alike items, hard tier: three measurement files, 622 items, 536 on the axis. No JevBench item and no JevBench file.",
10
+ "tier": "hard",
11
+ "map": "per_kind",
12
+ "dropped_leakage_flagged": 66
13
+ }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
  }