Pragya Β· test2
A 21.4M-parameter language model trained from scratch β for testing, not for production.
β οΈ Not a product
This is a scratch checkpoint used to shake out bugs in the training stack and sanity-check how the architecture scales. It is published because it might be useful as a small, from-scratch baseline β not because it is finished, aligned, or safe.
If you're looking for a capable small language model, this is not it.
Why this exists
A test artifact for the training pipeline. It was trained to answer questions like:
- Does the training loop converge cleanly at this architecture and scale?
- Do the fused Triton kernels produce numerically correct output vs. an eager reference?
- Does loss scale as expected with batch size, context length, and depth?
- Do the optimizer, LR schedule, and mixed-precision paths behave as designed?
The checkpoint itself is disposable. What it demonstrates is that the stack works end-to-end on a from-scratch 21M model.
Model card
| Parameters | 21.4M |
| Architecture | Decoder-only transformer |
| Layers | 16 |
| Attention | 12 query heads Β· 6 key/value heads (GQA) |
| Hidden size | 384 |
| Head dimension | 32 |
| FFN | SwiGLU |
| Context length | 1024 tokens |
| Vocabulary | 16,384 (byte-level BPE, tiktoken / RustBPE) |
| Positional | RoPE, base 10,000 |
| Attention mask | Dense causal |
| Embeddings | Tied (input embedding = LM head) |
| Training precision | BF16 |
On the training config: many flags in the internal config file β MoE, Mixture-of-Depths, MTP, quantization β are off or unused in this run. Treat the model card above as the spec, not the config file.
Training
| Unique tokens | ~2.9B |
| Epochs | 1 |
| Steps | 12,000 |
| Batch | 320 sequences Γ 1024 tokens |
| Optimizer | Muon (2D weights) + AdamW (embeddings, norms, biases) |
| LR schedule | 500-step warmup β WSD with 2,500-step decay |
| Final val perplexity | 16.0 |
Data mix
| Source | Role |
|---|---|
roneneldan/TinyStories |
Simple narrative English |
littlelearner/LittleCurriculum |
Structured educational text |
HuggingFaceTB/cosmopedia |
Synthetic textbook-style content |
HuggingFaceFW/fineweb-edu |
Filtered educational web text |
HuggingFaceTB/smollm-corpus |
Mixed high-quality web text |
English only. No additional filtering for toxicity, bias, or factual accuracy beyond each dataset's upstream filtering.
Capabilities
At 21M parameters, output drifts. That's the point of publishing it β you can see exactly where a from-scratch model at this size falls apart, and use it as a control when testing a new training change.
Sampling
For usable output, sample tightly:
temperature 0.6 β 0.9
top_k 20 β 50
top_p 0.9 (if using nucleus instead of top-k)
repetition_penalty 1.2 β 1.4
no_repeat_ngram 3
At this size, tight sampling matters more than usual β the model has very little capacity to recover from an off-distribution token. Loose settings turn grammatical drift into word salad within two sentences.
Files
pragya.pt β checkpoint (state_dict + config)
tokenizer/tokenizer.pkl β pickled tiktoken Encoding (vocab 16,384)
model.py β GPT architecture (GPT / GPTConfig)
inference.py β generate(), load_model(), sampling helpers
The tokenizer is a pickled
tiktoken.Encoding, not a HuggingFace tokenizer. Loading requires the model code from the same repo. It will not work withAutoTokenizer.from_pretrained(...).
Quick start
python inference.py --checkpoint pragya.pt --prompt "The Internet is" \
--top-k 20 --top-p 0.9 --temperature 0.6
inference.py handles checkpoint loading, KV-cached generation, sampling (top-k / top-p / repetition penalty / no-repeat n-grams), and streaming output. model.py contains the full architecture β no external dependencies beyond PyTorch and the tokenizer.
License
Apache 2.0. Each training dataset carries its own license β check upstream terms before redistributing derivatives.
Pragya Β· test2 Β· by Arush Kumar