compacttest-5m / README.md
Compactbot's picture
Add tokenizer_config.json, model.py, .gitattributes and README (renamed copy of gpt-s2.5-5m)
5f887f5 verified
|
Raw
History Blame
2.49 kB
metadata
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
  - tiny
  - tiny-lm
  - tiny-model
  - slm
  - small-language-model
  - from-scratch
  - gpt
  - gqa
  - swiglu
  - rope
  - rmsnorm
  - cpu-trained
library_name: transformers
metrics:
  - accuracy
model-index:
  - name: HellaSwag
    type: text-generation
    results: []

compacttest-5m

5.11M-parameter subword language model, trained from scratch on CPU.

This repo is a renamed copy of Compactbot/gpt-s2.5-5m, renamed at the request of @Datdanboi25 (see Compactbot/model-requests #2). Weights, config and architecture are identical to the original.

What it is

A small but genuine from-scratch causal LM in the GPT-X2.5 style:

  • Architecture: GQA (8 query / 2 key-value heads) + SwiGLU FFN + RoPE + RMSNorm, pre-norm, no biases, weight-tied input/output embeddings (the embedding matrix doubles as the LM head — no separate lm_head tensor).
  • Params: 5,114,112 (verified against the checkpoint).
  • Tokenizer: 8,192-vocab byte-level BPE (custom, not a standard HF tokenizer).
  • Trained: on CPU, ~94M TinyStories tokens, cosine LR schedule with warmup.

How to load

This is a self-contained nn.Module, not a transformers-native architecture.

import torch
from model import Model

m = Model()
from safetensors.torch import load_file
sd = load_file("model.safetensors")
m.load_state_dict(sd)
m.eval()

# greedy generation
with torch.no_grad():
    x = torch.tensor([[1]])
    for _ in range(60):
        logits = m(x[:, -512:])[:, -1, :]
        x = torch.cat([x, logits.argmax(-1, keepdim=True)], dim=1)

Honest scope

  • Greedy-coherent, sampling-fragile. At 5M params the model produces readable prose under greedy decoding but degrades noticeably under sampling. That is the expected behaviour at this scale, not a bug.
  • Intelligence index: 0.032 (see the original card for the full eval breakdown).
  • It is a reference build demonstrating that a 5M-param from-scratch subword LM is trainable and coherent on CPU. It is not a chat model.

Files

File Purpose
config.json Architecture config (matches model.py exactly)
model.py Self-contained architecture definition
model.safetensors Weights (5,114,112 params, F32)
tokenizer.json 8192-vocab BPE tokenizer
tokenizer_config.json Tokenizer metadata