rajpdus commited on
Commit
f3d2fa6
·
verified ·
1 Parent(s): cae8a43

Model card: add Requirements (transformers>=4.51) + fuller runnable example + VR sanity check

Browse files
Files changed (1) hide show
  1. README.md +27 -2
README.md CHANGED
@@ -22,16 +22,41 @@ for the [Tiny-ML Leaderboard](https://huggingface.co/spaces/Glint-Research/Tiny-
22
  [Jugnu](https://github.com/AltSlate-Labs/jugnu) family: the kept **R2** recipe (Qwen3 arch + **value residuals** +
23
  **Muon**) scaled to **25.2B tokens** under a **WSD** schedule with modest decay-phase educational upweighting.
24
 
 
 
 
 
 
 
 
 
 
25
  ## ⚠️ Load with `trust_remote_code=True`
26
 
27
  This model uses **value residuals** (a custom attention pathway: `v_i = v_proj_i(x) + λ_i·v0`). Stock
28
  `from_pretrained` would silently drop that pathway and degrade the model (~6 pts ARC-Easy, ~0.18 byte-ppl). It ships
29
- custom modeling code with `auto_map`, so load it VR-aware:
30
 
31
  ```python
32
  from transformers import AutoModelForCausalLM, AutoTokenizer
 
33
  tok = AutoTokenizer.from_pretrained("altslate/JugnuLM-110M-R2plus")
34
- model = AutoModelForCausalLM.from_pretrained("altslate/JugnuLM-110M-R2plus", trust_remote_code=True)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
35
  ```
36
 
37
  ## Results
 
22
  [Jugnu](https://github.com/AltSlate-Labs/jugnu) family: the kept **R2** recipe (Qwen3 arch + **value residuals** +
23
  **Muon**) scaled to **25.2B tokens** under a **WSD** schedule with modest decay-phase educational upweighting.
24
 
25
+ ## Requirements
26
+
27
+ ```bash
28
+ pip install "transformers>=4.51" torch safetensors
29
+ ```
30
+
31
+ `transformers>=4.51` is required (the model builds on the Qwen3 architecture). It's a standard
32
+ `AutoModelForCausalLM` otherwise — no extra packages.
33
+
34
  ## ⚠️ Load with `trust_remote_code=True`
35
 
36
  This model uses **value residuals** (a custom attention pathway: `v_i = v_proj_i(x) + λ_i·v0`). Stock
37
  `from_pretrained` would silently drop that pathway and degrade the model (~6 pts ARC-Easy, ~0.18 byte-ppl). It ships
38
+ custom modeling code with `auto_map`, so load it VR-aware (`trust_remote_code=True`):
39
 
40
  ```python
41
  from transformers import AutoModelForCausalLM, AutoTokenizer
42
+
43
  tok = AutoTokenizer.from_pretrained("altslate/JugnuLM-110M-R2plus")
44
+ model = AutoModelForCausalLM.from_pretrained(
45
+ "altslate/JugnuLM-110M-R2plus",
46
+ trust_remote_code=True, # required — rebuilds the value-residual pathway
47
+ ).eval()
48
+ # loads in fp32 by default; pass torch_dtype=torch.bfloat16 (transformers ≥5: dtype=...) to halve memory
49
+
50
+ ids = tok("The router will not connect to wifi, so I", return_tensors="pt").input_ids
51
+ out = model.generate(ids, max_new_tokens=40, do_sample=False)
52
+ print(tok.decode(out[0], skip_special_tokens=True))
53
+ ```
54
+
55
+ Sanity check that the value-residual pathway loaded (22 `vr_lambda` params, mean ≈ 0.48):
56
+
57
+ ```python
58
+ lam = [p.item() for n, p in model.named_parameters() if n.endswith("vr_lambda")]
59
+ assert len(lam) == 22, "value-residual pathway not loaded — did you pass trust_remote_code=True?"
60
  ```
61
 
62
  ## Results