ArkidMitra commited on
Commit
8b4cd7c
·
verified ·
1 Parent(s): 71d991f

Hopper (G) 1.3: new adapter weights (continued from 1.2); calibration map and serving unchanged; checksums and card

Browse files
Files changed (4) hide show
  1. CHECKSUMS.txt +2 -2
  2. README.md +38 -43
  3. adapter_config.json +10 -10
  4. adapter_model.safetensors +1 -1
CHECKSUMS.txt CHANGED
@@ -1,3 +1,3 @@
1
- 36d1214ea16648b4b7cf690c779ddd84e7a089589a7a62652a1c01332c24d3e1 adapter_config.json
2
- 8813cb2a1e44184bcdae017ace49cfc7cb00738c611423b6198c4b5ea553deaf adapter_model.safetensors
3
  217a2af396320c8a279e2f915fe17e67184cdf577668eb38147775184c259ade hopper.json
 
1
+ 40d70c507c1db76454ac0db030fed751105d9b08148a24fdd9946b4252f26e30 adapter_config.json
2
+ 2737fdc5c282393a84cbc4f38baf92de5ec2031a9375eeedf891818871613051 adapter_model.safetensors
3
  217a2af396320c8a279e2f915fe17e67184cdf577668eb38147775184c259ade hopper.json
README.md CHANGED
@@ -10,57 +10,52 @@ tags: [decision, classification, calibration, lora]
10
 
11
  A general-purpose version of [Hopper](https://huggingface.co/HopitAI/hopper): a LoRA adapter for Qwen3.5-4B that
12
  answers typed decision questions in one forward pass by reading the probability of each option letter, with a
13
- per-kind calibration map. Served with the Hopper code at https://github.com/hopit-ai/hopper (tag `g-1.2.0`).
14
 
15
- **Research and demo use only.** This adapter continues training from Hopper 1.0's adapter, whose training data
16
- included passages from RACE (non-commercial research only), and its training data also includes material made with
17
- LLM-based generation. Do not use it commercially.
18
 
19
- ## What it is
20
-
21
- - Base: Qwen/Qwen3.5-4B (revision `851bf6e`), LoRA rank 16, alpha 32, the same 12 modules as Hopper.
22
- - Training: continued from Hopper 1.0's adapter on Hopper's decision tasks (at maintenance doses), general-purpose
23
- sources (tabular record joins, CLINC150 intents, GSM8K arithmetic) and replay of public training data, with a fixed
24
- retention constraint against Hopper 1.0 on a held-out replay bank.
25
- - Serving: identical to Hopper 1.1.1, including the calibration map and the long-menu shortlist (more than 26
26
- options answered in two disclosed stages).
27
 
28
  ## Leaderboards (official)
29
 
30
  - **[Jev Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index)** (edition 0.2.1,
31
- 27 Sep 2026): **40.77, #16 of 68**, the highest of the 4B models (Decider 4B: 40.70). The row comes from a
32
- complete run of the suite that we scored ourselves with the Index kit, at the maintainer's request. The model
33
- outputs and scores are public at
34
- [`HopitAI/hopper-g-decision-index-results`](https://huggingface.co/datasets/HopitAI/hopper-g-decision-index-results).
35
- It replaced the Hopper 1.1.1 row (39.67).
36
- - **[JevBench](https://benchmarkheaven.com/jev-models)**: requested as a separate row
37
- ([issue #112](https://github.com/fstandhartinger/jevbench/issues/112)), not yet measured. Our own development check
38
- on held-out JevBench-style items (not an official score) put it level with Hopper 1.0 (+0.7 points, within noise),
39
- so we make no JevBench improvement claim. Hopper 1.0's official JevBench result is 59.43, #6 of 90 ranked
40
- (v1.4.2.1, 27 Sep 2026).
41
-
42
- ## Evaluation (our runs)
43
-
44
- On our local run of the Decision Index 0.2 suite (40 benchmarks, A10G, same serving code, only the adapter
45
- differing):
46
-
47
- | | Hopper 1.1.1 | Hopper (G) 1.2 |
48
- | --- | ---: | ---: |
49
- | balanced raw | 52.74 | 53.50 |
50
- | balanced skill | 37.10 | 38.07 |
51
- | GSM8K | 0.318 | 0.480 |
52
-
53
- Paired bootstrap of the balanced-raw difference: +0.76 (95 % interval +0.55 to +0.98). Seed 1 (trained independently) confirms: balanced raw 53.44 vs 52.74 (+0.70), with every Index area at or above Hopper 1.1.1.
54
- Not official scores. Hopper (G) is not tuned for JevBench; we make no claim there.
55
-
56
- ## Revisions
57
-
58
- - `d60a1d6`: the evaluated release (the revision given to the Decision Index and JevBench).
59
- - Later commit: `adapter_config.json` sets `"task_type": "CAUSAL_LM"` (it was `null`), for the Hub's metadata
60
- parser only. The weights, calibration map and outputs are unchanged.
61
 
62
  ## Limitations
63
 
64
  - English, 4B parameters; it reads options, it does not generate reasoning.
65
- - Calibration uses Hopper 1.1.1's map, fitted for Hopper 1.0's adapter; it has not been refitted for these weights.
 
66
  - Research and demo use only (see above).
 
 
 
 
 
 
10
 
11
  A general-purpose version of [Hopper](https://huggingface.co/HopitAI/hopper): a LoRA adapter for Qwen3.5-4B that
12
  answers typed decision questions in one forward pass by reading the probability of each option letter, with a
13
+ per-kind calibration map. Served with the Hopper code at https://github.com/hopit-ai/hopper.
14
 
15
+ **Current version: Hopper (G) 1.3** (tag `g-1.3.0`). Hopper (G) 1.2 stays available at
16
+ revision `d60a1d6` and is what the Decision Index row below measures.
 
17
 
18
+ **Research and demo use only.** These adapters continue training from Hopper 1.0's adapter, whose training data
19
+ included passages from RACE (non-commercial research only), and their training data also includes material made with
20
+ LLM-based generation. Do not use them commercially.
 
 
 
 
 
21
 
22
  ## Leaderboards (official)
23
 
24
  - **[Jev Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index)** (edition 0.2.1,
25
+ 27 Sep 2026): Hopper (G) 1.2 scores **40.77, #16 of 68**, the highest of the 4B models (Decider 4B: 40.70), from a
26
+ complete self-scored run ([results](https://huggingface.co/datasets/HopitAI/hopper-g-decision-index-results)).
27
+ Hopper (G) 1.3 has not been scored on the Index yet.
28
+ - **[JevBench](https://benchmarkheaven.com/jev-models)**: not yet measured. We have asked for Hopper (G) 1.3 to be
29
+ measured as a separate row ([issue #112](https://github.com/fstandhartinger/jevbench/issues/112)). Hopper 1.0's
30
+ official result is 59.43 (v1.4.2.2).
31
+
32
+ ## What changed in 1.3
33
+
34
+ - **Weights only.** Continued from Hopper (G) 1.2 on 6,400 new code-generated decision items (dated arithmetic, long
35
+ policies with exceptions and precedence, multi-table lookups, ambiguity, probability, safety and judging checklists,
36
+ trade-offs, paraphrase pairs, injected-instruction traps), every answer computed and checked by code, plus
37
+ maintenance and replay of earlier training data under the same retention constraint as 1.2.
38
+ - **Serving: unchanged.** Same code as Hopper 1.1.1 / Hopper (G) 1.2, same calibration map (byte-identical), eager
39
+ by default.
40
+
41
+ ## Evaluation (our runs, not official scores)
42
+
43
+ On our private held-out decision set (2,700 items, nine decision families, built for this purpose and never trained
44
+ on), 1.3 scores **+1.7 points over 1.2** (template bootstrap 95% interval +0.5 to +2.9). **This is below the +3.0 we
45
+ pre-registered as our own bar for this build**, and a smaller run with a quarter of the new data did about as well
46
+ (+2.1), so more of the same data did not help. We release it anyway as a disclosed, qualified release: on our
47
+ regression checks against Hopper 1.0 (document reading, numeric, paraphrase, abstention, routing, calibration and a
48
+ general-knowledge check) it passes every registered floor, with a pooled gain of +2.3 points (one-sided 95% lower
49
+ bound +1.9) and hard-tier calibration 83.5 under the shipped map. None of this predicts a JevBench result.
 
 
 
 
 
50
 
51
  ## Limitations
52
 
53
  - English, 4B parameters; it reads options, it does not generate reasoning.
54
+ - It rarely concludes "no match" or "cannot be determined" when an explicitly incomplete record leaves a multi-step
55
+ lookup unresolved; 1.3 did not fix this.
56
  - Research and demo use only (see above).
57
+
58
+ ## Revisions
59
+
60
+ - Tag `g-1.3.0`: Hopper (G) 1.3, the evaluated adapter bytes.
61
+ - `d60a1d6`: Hopper (G) 1.2 (the Decision Index row). `060b1bd`: its config-only fix.
adapter_config.json CHANGED
@@ -35,25 +35,25 @@
35
  "rank_pattern": {},
36
  "revision": null,
37
  "target_modules": [
 
 
 
 
38
  "in_proj_qkv",
39
  "q_proj",
40
- "gate_proj",
41
- "in_proj_z",
42
  "in_proj_a",
43
- "o_proj",
44
- "v_proj",
45
- "down_proj",
46
  "up_proj",
47
- "k_proj",
48
- "out_proj",
49
- "in_proj_b"
50
  ],
51
  "target_parameters": null,
52
- "task_type": "CAUSAL_LM",
53
  "trainable_token_indices": null,
54
  "use_bdlora": null,
55
  "use_dora": false,
56
  "use_qalora": false,
57
  "use_rslora": false,
58
  "velora_config": null
59
- }
 
35
  "rank_pattern": {},
36
  "revision": null,
37
  "target_modules": [
38
+ "v_proj",
39
+ "down_proj",
40
+ "out_proj",
41
+ "in_proj_z",
42
  "in_proj_qkv",
43
  "q_proj",
44
+ "k_proj",
 
45
  "in_proj_a",
 
 
 
46
  "up_proj",
47
+ "o_proj",
48
+ "in_proj_b",
49
+ "gate_proj"
50
  ],
51
  "target_parameters": null,
52
+ "task_type": null,
53
  "trainable_token_indices": null,
54
  "use_bdlora": null,
55
  "use_dora": false,
56
  "use_qalora": false,
57
  "use_rslora": false,
58
  "velora_config": null
59
+ }
adapter_model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:8813cb2a1e44184bcdae017ace49cfc7cb00738c611423b6198c4b5ea553deaf
3
  size 129927008
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2737fdc5c282393a84cbc4f38baf92de5ec2031a9375eeedf891818871613051
3
  size 129927008