Pragya · test3

A 27M-parameter decoder-only language model, trained from scratch — for testing, not for production.

Params Non-embed Layers Context Vocab Val PPL Val Acc License

Small language model · from-scratch training · custom Triton + PyTorch stack


⚠️ Not a product

This is a scratch checkpoint, not a finished language model. It exists to test a from-scratch training stack — kernels, optimizer, schedule, mixed precision — and to produce measurable, controlled improvements between runs. It is not aligned, factually reliable, or safe.

If you are looking for a capable small language model to build on, this is not it. If you are looking for a clean, well-documented baseline of what a 27M from-scratch model can do, keep reading.


TL;DR

Model

Backbone (non-embedding) 21.25M
MTP head 150K
Total parameters ~27M
Context 1024 tokens
Vocabulary 16,384 BPE
Layers 16
Heads 12 Q · 6 KV (GQA)

Results

Val perplexity 13.5
Val accuracy 49.7%
Val loss 2.603
CE loss 2.64
Unique tokens ~1.4B
Epochs 2

Delta over test2: −2.9 ppl · +6.6 accuracy points · 1,000 fewer steps.

Identical backbone (21.25M non-embedding) — only the training schedule, data budget, and MTP head changed.

The interesting part of this release is not the model — it's that a controlled change to the training schedule and data budget produced a predicted, measurable improvement at fixed compute, with no change to the model itself. See Progression.


Why this exists

Pragya is a from-scratch small language model project. Every layer, kernel, and training loop is hand-written; nothing is imported from a pretrained checkpoint. test3 is the third checkpoint from that pipeline and it exists to answer one class of question:

Does the training stack actually work, and can I move the outcome with intent?

Concretely, this run was built to verify:

  • Numerical correctness — fused Triton kernels (attention post-processing, QK-norm + RoPE, add + RMSNorm, MoE group-GEMMs) match an eager reference to bf16 tolerance.
  • Convergence stability — no loss spikes, no diverging runs, no NaN gradients across 11,000 steps.
  • Schedule sensitivity — a longer Warmup–Stable–Decay (WSD) tail produces a sharp, measurable drop in loss.
  • Data-budget sensitivity — fewer unique tokens seen twice beats more unique tokens seen once, at fixed compute.

The model itself is disposable. The demonstration is not.


Model card

Backbone (non-embedding) 21.25M
MTP head 150K
Total parameters ~27M
Architecture Decoder-only transformer
Layers 16
Attention 12 query heads · 6 key/value heads (GQA)
Hidden size 384
Head dimension 32
FFN SwiGLU (ffn_mult = 8/3)
Context length 1024 tokens
Vocabulary 16,384 — byte-level BPE (tiktoken / RustBPE)
Positional encoding RoPE, base 10,000
Attention mask Dense causal
Embeddings Tied (input embedding = LM head)
MTP auxiliary head Enabled · depth 1 · λ = 0.1
Training precision BF16

On the training config file: many flags in the internal config.py — MoE, Mixture-of-Depths, quantization — are off or unused in this run. Treat the model card above as the spec. The config file is shared across experiments and does not describe this checkpoint.


Training

Compute

Unique tokens ~1.4B
Epochs 2
Steps 11,000
Batch 320 × 1024
Precision BF16

Recipe

Optimizer Muon + AdamW
Schedule Linear warmup → WSD
Warmup 500 steps
Stable 8,000 steps
Decay 2,500 steps (cosine)

Final metrics: CE loss 2.64 · Val loss 2.603 · Val perplexity 13.5 · Val accuracy 49.7%

Data mix

All English. No additional filtering for toxicity, bias, or factual accuracy beyond each dataset's own upstream filters.

Source Role
roneneldan/TinyStories Simple narrative English
littlelearner/LittleCurriculum Structured educational text
HuggingFaceTB/cosmopedia Synthetic textbook-style content
HuggingFaceFW/fineweb-edu Filtered educational web text
HuggingFaceTB/smollm-corpus Mixed high-quality web text

Progression

Two checkpoints from the same training stack. Same model, same compute budget per run, different configurations.

test2 test3
Backbone (non-embedding) 21.25M 21.25M
MTP head off on (+150K)
Total parameters ~27M ~27M
Unique tokens ~2.9B ~1.4B
Epochs 1 2
Steps 12,000 11,000
Decay steps 700 2,500
Context 1024 1024
Val loss 2.797 2.603
Val perplexity 16.4 13.5
Val accuracy 43.1% 49.7%

Delta (test2 → test3): −0.194 nats · −2.9 ppl · +6.6 accuracy points, at 1,000 fewer steps, with no change to the model backbone.

Difference table

What changed between test2 and test3, at a glance:

test2 test3 Δ Direction
Backbone (non-embedding) 21.25M 21.25M 0 = identical
MTP head off on +150K ✚ added
Total parameters ~27M ~27M ≈ 0 = same
Unique tokens ~2.9B ~1.4B −1.5B ↓ smaller
Epochs 1 2 +1 ↑ more
Tokens seen ~2.9B ~2.8B −0.1B ≈ same
Steps 12,000 11,000 −1,000 ↓ fewer
Decay steps 700 2,500 +1,800 ↑ 3.6×
Decay fraction 5.8% 22.7% +16.9 pp ↑ longer tail
Context 1024 1024 0 = same
Val loss 2.797 2.603 −0.194 ↓ better
Val perplexity 16.4 13.5 −2.9 ↓ better
Val accuracy 43.1% 49.7% +6.6 pp ↑ better

Read this table as: identical model, same compute budget, two changes to the training recipe, results improved across every metric, and the run that won was 1,000 steps shorter.

Two rows are worth pausing on:

  • Tokens seen is essentially flat (~2.8B vs ~2.9B). The improvement is not from showing the model more data — it's from showing it the same amount of data in a different arrangement.
  • Decay fraction jumped 4×. Everything else is either neutral (context, backbone) or a smaller factor (MTP +0.5% params). If you had to pick one line as the probable cause, it's this one.

What actually moved

Two changes between test2 and test3:

  1. Longer WSD decay. 700 → 2,500 decay steps (5.8% → 22.7% of total steps).
  2. Data schedule. 2.9B unique × 1 epoch → 1.4B unique × 2 epochs. Same tokens-seen budget.

Plus one minor addition:

  1. MTP head. Added a 150K-parameter Multi-Token Prediction auxiliary (+0.5% total params). The backbone itself is unchanged (21.25M non-embedding in both).

The observed loss curve is dominated by the decay phase — the plateau was contributing ~0.2 ppl per 1,000 steps, and the decay contributed ~1.0 ppl per 1,000 steps. The backbone is unchanged between the two runs, so the improvement comes from the training schedule, the data arrangement, or the MTP head — not from a bigger model. All three changed at once, so the attribution is suggestive, not isolated.

Both metrics moved in lockstep — ppl −17.7% relative, accuracy +15.3% relative. If only perplexity had improved, the win could have been a confidence effect; accuracy rising alongside it confirms the model is actually more correct, not just more certain.

Open questions

  • Which lever did the work? Schedule, data arrangement, and MTP all changed together. A controlled A/B would isolate the contribution of each.
  • Is 22.7% decay the right fraction? Test3's decay was still dropping when it ended — headroom remains. A longer decay on the same total step budget is worth testing.
  • Does MTP help at this scale? The 150K MTP parameters are ~0.5% of the model. Whether the auxiliary loss buys anything measurable at 21.25M non-embedding is unresolved.

These are the next runs, not claims about this checkpoint.


Capabilities

✅ It can

  • Continue English prompts for a sentence or two
  • Produce grammatically plausible text in a few registers (narrative, encyclopedic, instructional)
  • Serve as a fine-tuning starting point
  • Act as a control baseline for training experiments
  • Run on CPU or a single consumer GPU

❌ It can't

  • Answer questions factually
  • Follow instructions
  • Stay coherent past the first paragraph
  • Handle non-English input
  • Deal with context longer than 1024 tokens

At 21.25M non-embedding parameters, output drifts. That's the point of publishing it — you can see exactly where a from-scratch model at this size falls apart, and use it as a control when testing a new training change.


Sampling

For usable output, sample tightly:

temperature        0.6 – 0.9
top_k              20 – 50
top_p              0.9        (if using nucleus instead of top-k)
repetition_penalty 1.2 – 1.4
no_repeat_ngram    3

At this size, tight sampling matters more than usual — the model has very little capacity to recover from an off-distribution token. Loose settings turn grammatical drift into word salad within two sentences.


Quick start

Files

pragya.pt                 ← checkpoint (state_dict + config)
tokenizer/tokenizer.pkl   ← pickled tiktoken Encoding (vocab 16,384)
model.py                  ← GPT architecture (GPT / GPTConfig)
inference.py              ← generate(), load_model(), sampling helpers

⚠️ The tokenizer is a pickled tiktoken.Encoding, not a HuggingFace tokenizer. Loading requires the model code from the same repo. It will not work with AutoTokenizer.from_pretrained(...).

CLI

python inference.py \
    --checkpoint pragya.pt \
    --prompt "The Internet is" \
    --top-k 20 \
    --top-p 0.9 \
    --temperature 0.6 \
    --max-new-tokens 200

Programmatic

from inference import load_model, generate, pick_device
from tokenizer import TokenizerWrapper

device    = pick_device(None)
tokenizer = TokenizerWrapper("./tokenizer")
model, _  = load_model("pragya.pt", device, tokenizer)

texts = generate(
    model, tokenizer,
    prompt="The Internet is",
    device=device,
    max_new_tokens=200,
    temperature=0.6,
    top_k=20,
    top_p=0.9,
    repetition_penalty=1.3,
    no_repeat_ngram_size=3,
)
print(texts[0])

inference.py handles checkpoint loading, KV-cached generation, the sampling stack (top-k / top-p / repetition penalty / no-repeat n-grams), and streaming output.

model.py contains the full architecture — no external dependencies beyond PyTorch and the tokenizer.


Limitations

  • Small scale. Output is grammatical for a sentence or two and then drifts into invented facts or register changes. This is a property of the parameter count, not a bug.
  • Mixed training corpus. The model blends encyclopedic, narrative, and instructional registers and sometimes blends them mid-paragraph.
  • English only. Byte-level BPE encodes other languages, but the model has no training signal for them.
  • No instruction tuning. Any instruction-following ability is incidental. For instruction behavior, fine-tune on a supervised dataset first.
  • No alignment or safety filtering. The model can and will produce biased, false, or toxic continuations.

FAQ

Can I load this with AutoModelForCausalLM.from_pretrained?

No. The tokenizer is a pickled tiktoken.Encoding and the model uses custom code in model.py. Loading requires the model and inference scripts from this repo.

Does this follow instructions?

No. It's a base pretraining checkpoint, not an instruction-tuned model. It continues text; it doesn't answer questions. Fine-tune it on a supervised dataset if you want instruction behavior.

Why is the val accuracy 49.7% and perplexity 13.5?

Because the vocabulary is 16,384 tokens. Random next-token accuracy on a 16k vocab is ~0.006%. Going from 43.1% (test2) to 49.7% (test3) means the model is now predicting the correct next token on roughly half of all positions. That is a property of the training data — formulaic narrative + textbook text — as much as the model. Real generalization is much weaker than this number suggests.

Why is the training corpus so small?

Because this is a test run at fixed compute, not a production model. 1.4B unique tokens is enough to see whether the pipeline works and whether schedule changes matter. It's not enough to build a model that knows things.

Can I use this commercially?

The model weights are Apache 2.0. Each training dataset carries its own license — check upstream terms before redistributing derivatives. fineweb-edu, TinyStories, and smollm-corpus all have permissive licenses; verify the current terms before commercial use.


Citation

If you use this checkpoint in research, please link to this repository. There is no paper.

@misc{pragya-test3,
  title  = {Pragya test3: a 27M from-scratch decoder-only language model},
  author = {Arush Kumar},
  year   = {2026},
  note   = {Test checkpoint, not for production use},
  howpublished = {\url{https://huggingface.co/ArushKumar/pragya-test3}}
}

License

Apache 2.0. Each training dataset carries its own license — check upstream terms before redistributing derivatives.

Pragya · test3

trained from scratch · built on a custom Triton + PyTorch stack · for testing, not production
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train RuVI-AI/pragya-test-3