Mamb2_8B_Recall / docs /RESULTS.md
EndlessChasing's picture
Publish verified Mamba2-8B Resurface adapter and reproducibility package
5b7b27a verified
|
Raw History Blame Contribute Delete
3.62 kB

Mamba2-8B source Resurface control: completed results

This is a post-D Resurface-style adapter on the uncompressed, pure NVIDIA Mamba2-8B source. The official BF16 checkpoint is cast to FP16 in the native Mamba2 runtime. All 8,236,999,680 base parameters stay frozen. The serialized adapter contains 1,154,104 FP16 parameters across 56 layers and is 2,374,143 bytes.

Matched four-arm comparison

All four arms were checked against the same 768 tokenized MK prompts and the same 130 WikiText-2 validation windows (264,764 next-token targets). For MK, the denominator is 384 normal prompts and 384 target-removed controls per arm.

Arm Normal MK Target removed WikiText-2 PPL
Source FP16 147/384 (38.28%) 0/384 7.33418
Source FP16 + Resurface 365/384 (95.05%) 0/384 7.05206
E8/W5 compressed, readapted base 91/384 (23.70%) 0/384 7.62240
E8/W5 compressed, readapted base + Resurface 340/384 (88.54%) 0/384 7.59316

On the source, the adapter gained 222 previously wrong normal prompts and lost 4 previously correct prompts: net +218/384, or +56.77 percentage points. Its PPL changed by -0.28211 (-3.85%). The compressed arm's normal recall gain was +249/384 (+64.84 points), with a PPL change of -0.02923 (-0.38%). The 8.07-point difference between these gains is descriptive; it does not isolate the effect of quantization, since the compressed base also had 448 small-tensor readaptation updates and the two adapters used different frozen base teachers for their KL losses.

Run and verification

  • Final candidate: 1,536 successful optimizer updates, 1,542 attempts, six numerical overflow retries within the frozen eight-retry budget. All 1,536 TRAIN examples appeared once; no checkpoint was chosen using validation.
  • Four saved optimizer checkpoints independently matched their recorded SHA-256 values. The final adapter passed exact FP16 tensor roundtrip and source-checkpoint binding checks. Adapter SHA-256: e8b2b4dfe69f8e85dc9e147c9aeaa3297cff14f4043e558bb1795ad476c1fca0.
  • Evaluation loaded the serialized adapter, scored source and adapter in one process, then removed it. The first eight baseline MK outputs were reproduced exactly after removal. Frozen-base parameter identity, version counters and absence of gradients were checked; a full bytewise rescan of all 16 GB of base weights was not performed.
  • compare_four_arms.py required exact ordered prompt token hashes for all 768 MK rows and exact token hashes for every PPL window across all four arms. Both checks passed. The old compressed reports were also verified by their pinned SHA-256 values.
  • Peak PyTorch CUDA allocation was 34.59 GB during two-model training and 17.08 GB during single-model paired evaluation. These are GPU allocation measurements, not total device or host memory requirements.

The protocol was frozen before this adapter was trained. Raw training, paired evaluation and four-arm comparison reports are included for audit. The numeric CONFIRM template family and instances had already been observed in the earlier compressed project, and WikiText-2 validation was used in prior development. These results establish performance on this specified replay protocol; they do not establish generalization to unseen template families, longer recall distances, or untouched prose.