Explore how a balanced token-ID hash selects two narrow residual experts while a shared dense SwiGLU path preserves contextual computation.
| Measurement | Initial | Final |
|---|---|---|
| Pretraining held-out NLL | 4.8554 @ 1k | 2.6615 @ 76k |
| Pretraining held-out PPL | 128.43 | 14.32 |
| Matched SFT NLL | 3.4805 | 2.9669 |
| Natural-gold SFT NLL | 3.0644 | 2.6866 |
| Training coverage | 20B pretrain | 31.37M supervised |
Pretraining evaluation is held out. SFT diagnostics use a 672-example matched set and a 28-example natural-gold set; the latter is too small for a stable capability claim.
The model response will stream here token by token.
Chat endpoint · temperature 0.7 · top-k 40 · top-p 0.9 · repetition penalty 1.1 · maximum 128 new tokens.
This is a qualitative demonstration of the released SFT checkpoint, not an evaluation result.