COMPLEXITY TR-HASH-0.5B companion
Final research checkpoint

TR-HASH-0.5B

Explore how a balanced token-ID hash selects two narrow residual experts while a shared dense SwiGLU path preserves contextual computation.

492.1M parameters 20B pretraining tokens 24 layers · fixed top-2
Interactive companion — not additional experimental evidence.
Token
Layer-specific deterministic routes
Checkpoint-derived tokenizer IDs and fixed top-2 assignments.
E0E1E2E3 Layer 10.50.5 Layer 60.50.5 Layer 120.50.5 Layer 180.50.5 Layer 240.50.5
Token “routing” follows the persisted route table of the final checkpoint. Every occurrence keeps the same expert pair at a given layer, while its hidden-state input still changes with context.
Recorded training diagnostics
Initial and final scheduled evaluations · NLL, lower is better
Initial Final
MeasurementInitialFinal
Pretraining held-out NLL4.8554 @ 1k2.6615 @ 76k
Pretraining held-out PPL128.4314.32
Matched SFT NLL3.48052.9669
Natural-gold SFT NLL3.06442.6866
Training coverage20B pretrain31.37M supervised

Pretraining evaluation is held out. SFT diagnostics use a 672-example matched set and a 28-example natural-gold set; the latter is too small for a stable capability claim.

TR-HASH-0.5B ready

The model response will stream here token by token.

Decode configuration

Chat endpoint · temperature 0.7 · top-k 40 · top-p 0.9 · repetition penalty 1.1 · maximum 128 new tokens.

This is a qualitative demonstration of the released SFT checkpoint, not an evaluation result.

TR-HASH-0.5B · live qualitative generation