Mamb2_8B_Recall / docs /PROTOCOL.md
EndlessChasing's picture
Publish verified Mamba2-8B Resurface adapter and reproducibility package
5b7b27a verified
|
Raw History Blame Contribute Delete
4.54 kB
# Full precision Mamba2-8B Resurface control
Declared 2026-09-28 before fitting this repository's adapter. The purpose is to
measure the effect of the same readout adapter and training budget on NVIDIA's
uncompressed pure Mamba2-8B. The E8/W5 experiment is a historical comparator,
not a training dependency. All metrics are reported at the actually measured
precision and protocol, without assuming equality to native Megatron BF16.
## Frozen model and intervention
Load `nvidia/mamba2-8b-3t-4k` revision
`b915550c63ba9359f88f44d1f6a600d85af27302`, checkpoint SHA256
`47c2766f6aad89d73beafbeaecb334aab902d7370906d081764a90bb7a8bbbcb`
and tokenizer SHA256
`5862e2f71caf762bc9845662be5fec2867deb58d874568235a02a36c5111cd09`.
Cast the 507 original BF16 tensors to FP16 in the native state-spaces Mamba2
runtime; freeze all 8,236,999,680 base parameters. Verify strict source tensor
mapping, model geometry and no shared embedding/output storage.
Use the same post-D adapter as the compressed experiment at all 56 native
gated RMSNorm inputs: `y + sigmoid(w dot u+b) * g*(V@y)` across 128 heads.
Its 224 tensors contain 1,154,104 FP32 trainable master parameters, exported
as FP16 for inference. Initialize V=0, g=1, w=0, b=-4; use soft sigmoid at
train and eval, no task switch, EMA, or new recurrent cache. The adapter is
external to the frozen base; check its identities and absence of base grads.
## Training
Reuse the *same exact data generator contract and seeds* as the compressed
adapter: 1,536 TRAIN numeric bindings over three templates and N=16/64,
disjoint six-digit key/value intervals; full 256K CE on every answer suffix
token. Use seed 2026092803 and `torch.randperm(1536)` once for example order.
Each step pairs one MK example with a 512-token segment from the already
prepared 448 WikiText-2 TRAIN windows. Their historical manifest and file
SHA256 are pinned by the train command. Heldout and validation windows are
not included in training. Reusing this fixed TRAIN text gives the control the
same training exposures as the compressed arm.
Use a second, independently loaded **uncompressed FP16** model as the frozen
prose teacher. This is the direct counterpart to the compressed experiment's
teacher, which was its own unadapted compressed base. The fixed per-step loss is
`MK answer CE + 0.5 prose CE + 0.5 KL(own unadapted teacher || student) + 3 C`,
where C is the same prose-only router closure, with budget 0.006 and excess
coefficient 10. Temperature 1, 511 prose targets per step, 64-token staged
head chunks, and one optimizer update after both task gradients.
AdamW FP32 masters: V/g LR 1e-4, w/b LR 3e-4, betas (.9,.999), eps 1e-8,
weight decay 0, clip norm 1. Multiplicative schedule
`0.1+0.9*(1+cos(pi*j/1535))/2` at successful index j. Native FP16 forward,
block checkpointing, gradient scale 1024 with growth interval 2000. Overflow
retries the exact same pair without optimizer/master update, at most 8 retries
and 1,544 total attempts. Run 1,536 **successful** updates; the final step is
the only candidate. Small GPU smokes are discarded before formal training.
## Evaluation and interpretation
After training, restore the actual serialized FP16 adapter onto the same
uncompressed base. Evaluate baseline, adapter enabled and restored baseline
in a single process, with identical tokenizer, MK prompts and full WikiText-2
validation windows. For MK use full 256K greedy generation up to 12 tokens,
first standalone six-digit number, normal and target-removed controls, native
prefill plus recurrent decode with fresh FP16 cache. For PPL use all 130
nonoverlapping reset windows and all 264,764 next-token targets. Record all
raw scores, differences and file hashes; no early selection or task switch.
The earlier compressed arm's independent CONFIRM set has now been observed and
uses the same three template families. Running this source control on it gives
a matched historical comparison, **not a newly untouched holdout**. The
WikiText-2 validation text also informed earlier development. We will not
promote a source+adapter result to a general recall conclusion without new
templates, longer distances and a truly untouched corpus.
Report all four observed arms explicitly: source FP16, source+adapter,
compressed/readapted, compressed/readapted+adapter. The source and compressed
teachers differ by construction. Report PPL and recall together; any cross-arm
claim remains limited by shared test protocols and the historical numerical
replay discrepancy in the compressed project.