Instructions to use HopitAI/hopper-g with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use HopitAI/hopper-g with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B") model = PeftModel.from_pretrained(base_model, "HopitAI/hopper-g") - Notebooks
- Google Colab
- Kaggle
Hopper (G) 1.2: adapter, calibration map, checksums and model card
Browse files- CHECKSUMS.txt +3 -0
- README.md +46 -0
- adapter_config.json +59 -0
- adapter_model.safetensors +3 -0
- hopper.json +14 -0
CHECKSUMS.txt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
5e6c17beae26839e65be9805023c46397a073af83edad87c29b111d5d8f1c490 adapter_config.json
|
| 2 |
+
8813cb2a1e44184bcdae017ace49cfc7cb00738c611423b6198c4b5ea553deaf adapter_model.safetensors
|
| 3 |
+
217a2af396320c8a279e2f915fe17e67184cdf577668eb38147775184c259ade hopper.json
|
README.md
ADDED
|
@@ -0,0 +1,46 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: research-and-demo
|
| 4 |
+
base_model: Qwen/Qwen3.5-4B
|
| 5 |
+
library_name: peft
|
| 6 |
+
tags: [decision, classification, calibration, lora]
|
| 7 |
+
---
|
| 8 |
+
|
| 9 |
+
# Hopper (G)
|
| 10 |
+
|
| 11 |
+
A general-purpose version of [Hopper](https://huggingface.co/HopitAI/hopper): a LoRA adapter for Qwen3.5-4B that
|
| 12 |
+
answers typed decision questions in one forward pass by reading the probability of each option letter, with a
|
| 13 |
+
per-kind calibration map. Served with the Hopper code at https://github.com/hopit-ai/hopper (tag `g-1.2.0`).
|
| 14 |
+
|
| 15 |
+
**Research and demo use only.** This adapter continues training from Hopper 1.0's adapter, whose training data
|
| 16 |
+
included passages from RACE (non-commercial research only), and its training data also includes material made with
|
| 17 |
+
LLM-based generation. Do not use it commercially.
|
| 18 |
+
|
| 19 |
+
## What it is
|
| 20 |
+
|
| 21 |
+
- Base: Qwen/Qwen3.5-4B (revision `851bf6e`), LoRA rank 16, alpha 32, the same 12 modules as Hopper.
|
| 22 |
+
- Training: continued from Hopper 1.0's adapter on Hopper's decision tasks (at maintenance doses), general-purpose
|
| 23 |
+
sources (tabular record joins, CLINC150 intents, GSM8K arithmetic) and replay of public training data, with a fixed
|
| 24 |
+
retention constraint against Hopper 1.0 on a held-out replay bank.
|
| 25 |
+
- Serving: identical to Hopper 1.1.1, including the calibration map and the long-menu shortlist (more than 26
|
| 26 |
+
options answered in two disclosed stages).
|
| 27 |
+
|
| 28 |
+
## Evaluation (our runs)
|
| 29 |
+
|
| 30 |
+
On our local run of the Decision Index 0.2 suite (40 benchmarks, A10G, same serving code, only the adapter
|
| 31 |
+
differing):
|
| 32 |
+
|
| 33 |
+
| | Hopper 1.1.1 | Hopper (G) 1.2 |
|
| 34 |
+
| --- | ---: | ---: |
|
| 35 |
+
| balanced raw | 52.74 | 53.50 |
|
| 36 |
+
| balanced skill | 37.10 | 38.07 |
|
| 37 |
+
| GSM8K | 0.318 | 0.480 |
|
| 38 |
+
|
| 39 |
+
Paired bootstrap of the balanced-raw difference: +0.76 (95 % interval +0.55 to +0.98). Seed 1 (trained independently) confirms: balanced raw 53.44 vs 52.74 (+0.70), with every Index area at or above Hopper 1.1.1.
|
| 40 |
+
Not official scores. Hopper (G) is not tuned for JevBench; we make no claim there.
|
| 41 |
+
|
| 42 |
+
## Limitations
|
| 43 |
+
|
| 44 |
+
- English, 4B parameters; it reads options, it does not generate reasoning.
|
| 45 |
+
- Calibration uses Hopper 1.1.1's map, fitted for Hopper 1.0's adapter; it has not been refitted for these weights.
|
| 46 |
+
- Research and demo use only (see above).
|
adapter_config.json
ADDED
|
@@ -0,0 +1,59 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": {
|
| 6 |
+
"base_model_class": "Qwen3_5ForCausalLM",
|
| 7 |
+
"parent_library": "transformers.models.qwen3_5.modeling_qwen3_5"
|
| 8 |
+
},
|
| 9 |
+
"base_model_name_or_path": "Qwen/Qwen3.5-4B",
|
| 10 |
+
"bias": "none",
|
| 11 |
+
"corda_config": null,
|
| 12 |
+
"ensure_weight_tying": false,
|
| 13 |
+
"eva_config": null,
|
| 14 |
+
"exclude_modules": null,
|
| 15 |
+
"fan_in_fan_out": false,
|
| 16 |
+
"inference_mode": true,
|
| 17 |
+
"init_lora_weights": true,
|
| 18 |
+
"kasa_config": null,
|
| 19 |
+
"layer_replication": null,
|
| 20 |
+
"layers_pattern": null,
|
| 21 |
+
"layers_to_transform": null,
|
| 22 |
+
"loftq_config": {},
|
| 23 |
+
"lora_alpha": 32,
|
| 24 |
+
"lora_bias": false,
|
| 25 |
+
"lora_dropout": 0.05,
|
| 26 |
+
"lora_ga_config": null,
|
| 27 |
+
"megatron_config": null,
|
| 28 |
+
"megatron_core": "megatron.core",
|
| 29 |
+
"modules_to_save": null,
|
| 30 |
+
"monteclora_config": null,
|
| 31 |
+
"peft_type": "LORA",
|
| 32 |
+
"peft_version": "0.21.0",
|
| 33 |
+
"qalora_group_size": 16,
|
| 34 |
+
"r": 16,
|
| 35 |
+
"rank_pattern": {},
|
| 36 |
+
"revision": null,
|
| 37 |
+
"target_modules": [
|
| 38 |
+
"in_proj_qkv",
|
| 39 |
+
"q_proj",
|
| 40 |
+
"gate_proj",
|
| 41 |
+
"in_proj_z",
|
| 42 |
+
"in_proj_a",
|
| 43 |
+
"o_proj",
|
| 44 |
+
"v_proj",
|
| 45 |
+
"down_proj",
|
| 46 |
+
"up_proj",
|
| 47 |
+
"k_proj",
|
| 48 |
+
"out_proj",
|
| 49 |
+
"in_proj_b"
|
| 50 |
+
],
|
| 51 |
+
"target_parameters": null,
|
| 52 |
+
"task_type": null,
|
| 53 |
+
"trainable_token_indices": null,
|
| 54 |
+
"use_bdlora": null,
|
| 55 |
+
"use_dora": false,
|
| 56 |
+
"use_qalora": false,
|
| 57 |
+
"use_rslora": false,
|
| 58 |
+
"velora_config": null
|
| 59 |
+
}
|
adapter_model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:8813cb2a1e44184bcdae017ace49cfc7cb00738c611423b6198c4b5ea553deaf
|
| 3 |
+
size 129927008
|
hopper.json
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"kind": "per_kind",
|
| 3 |
+
"temperatures": {
|
| 4 |
+
"choice": 0.7898505975532796,
|
| 5 |
+
"noul": 0.7531284680253558,
|
| 6 |
+
"score": 0.8997033695617845
|
| 7 |
+
},
|
| 8 |
+
"fitted_on": {
|
| 9 |
+
"description": "our own held-out JevBench-style look-alike items, hard tier: three measurement files, 622 items, 536 on the axis. No JevBench item and no JevBench file.",
|
| 10 |
+
"tier": "hard",
|
| 11 |
+
"map": "per_kind",
|
| 12 |
+
"dropped_leakage_flagged": 66
|
| 13 |
+
}
|
| 14 |
+
}
|