|
Download docs/PROTOCOL.md from EndlessChasing/Mamb2_8B_Recall: direct link, hf CLI and curl.
- Browser
- Download file 4.54 kB
-
https://huggingface.co/EndlessChasing/Mamb2_8B_Recall/resolve/main/docs/PROTOCOL.md
- Command line
-
hf download hf://EndlessChasing/Mamb2_8B_Recall/docs/PROTOCOL.md
-
curl -L -o PROTOCOL.md https://huggingface.co/EndlessChasing/Mamb2_8B_Recall/resolve/main/docs/PROTOCOL.md
4.54 kB
| # Full precision Mamba2-8B Resurface control | |
| Declared 2026-09-28 before fitting this repository's adapter. The purpose is to | |
| measure the effect of the same readout adapter and training budget on NVIDIA's | |
| uncompressed pure Mamba2-8B. The E8/W5 experiment is a historical comparator, | |
| not a training dependency. All metrics are reported at the actually measured | |
| precision and protocol, without assuming equality to native Megatron BF16. | |
| ## Frozen model and intervention | |
| Load `nvidia/mamba2-8b-3t-4k` revision | |
| `b915550c63ba9359f88f44d1f6a600d85af27302`, checkpoint SHA256 | |
| `47c2766f6aad89d73beafbeaecb334aab902d7370906d081764a90bb7a8bbbcb` | |
| and tokenizer SHA256 | |
| `5862e2f71caf762bc9845662be5fec2867deb58d874568235a02a36c5111cd09`. | |
| Cast the 507 original BF16 tensors to FP16 in the native state-spaces Mamba2 | |
| runtime; freeze all 8,236,999,680 base parameters. Verify strict source tensor | |
| mapping, model geometry and no shared embedding/output storage. | |
| Use the same post-D adapter as the compressed experiment at all 56 native | |
| gated RMSNorm inputs: `y + sigmoid(w dot u+b) * g*(V@y)` across 128 heads. | |
| Its 224 tensors contain 1,154,104 FP32 trainable master parameters, exported | |
| as FP16 for inference. Initialize V=0, g=1, w=0, b=-4; use soft sigmoid at | |
| train and eval, no task switch, EMA, or new recurrent cache. The adapter is | |
| external to the frozen base; check its identities and absence of base grads. | |
| ## Training | |
| Reuse the *same exact data generator contract and seeds* as the compressed | |
| adapter: 1,536 TRAIN numeric bindings over three templates and N=16/64, | |
| disjoint six-digit key/value intervals; full 256K CE on every answer suffix | |
| token. Use seed 2026092803 and `torch.randperm(1536)` once for example order. | |
| Each step pairs one MK example with a 512-token segment from the already | |
| prepared 448 WikiText-2 TRAIN windows. Their historical manifest and file | |
| SHA256 are pinned by the train command. Heldout and validation windows are | |
| not included in training. Reusing this fixed TRAIN text gives the control the | |
| same training exposures as the compressed arm. | |
| Use a second, independently loaded **uncompressed FP16** model as the frozen | |
| prose teacher. This is the direct counterpart to the compressed experiment's | |
| teacher, which was its own unadapted compressed base. The fixed per-step loss is | |
| `MK answer CE + 0.5 prose CE + 0.5 KL(own unadapted teacher || student) + 3 C`, | |
| where C is the same prose-only router closure, with budget 0.006 and excess | |
| coefficient 10. Temperature 1, 511 prose targets per step, 64-token staged | |
| head chunks, and one optimizer update after both task gradients. | |
| AdamW FP32 masters: V/g LR 1e-4, w/b LR 3e-4, betas (.9,.999), eps 1e-8, | |
| weight decay 0, clip norm 1. Multiplicative schedule | |
| `0.1+0.9*(1+cos(pi*j/1535))/2` at successful index j. Native FP16 forward, | |
| block checkpointing, gradient scale 1024 with growth interval 2000. Overflow | |
| retries the exact same pair without optimizer/master update, at most 8 retries | |
| and 1,544 total attempts. Run 1,536 **successful** updates; the final step is | |
| the only candidate. Small GPU smokes are discarded before formal training. | |
| ## Evaluation and interpretation | |
| After training, restore the actual serialized FP16 adapter onto the same | |
| uncompressed base. Evaluate baseline, adapter enabled and restored baseline | |
| in a single process, with identical tokenizer, MK prompts and full WikiText-2 | |
| validation windows. For MK use full 256K greedy generation up to 12 tokens, | |
| first standalone six-digit number, normal and target-removed controls, native | |
| prefill plus recurrent decode with fresh FP16 cache. For PPL use all 130 | |
| nonoverlapping reset windows and all 264,764 next-token targets. Record all | |
| raw scores, differences and file hashes; no early selection or task switch. | |
| The earlier compressed arm's independent CONFIRM set has now been observed and | |
| uses the same three template families. Running this source control on it gives | |
| a matched historical comparison, **not a newly untouched holdout**. The | |
| WikiText-2 validation text also informed earlier development. We will not | |
| promote a source+adapter result to a general recall conclusion without new | |
| templates, longer distances and a truly untouched corpus. | |
| Report all four observed arms explicitly: source FP16, source+adapter, | |
| compressed/readapted, compressed/readapted+adapter. The source and compressed | |
| teachers differ by construction. Report PPL and recall together; any cross-arm | |
| claim remains limited by shared test protocols and the historical numerical | |
| replay discrepancy in the compressed project. | |