compacttest-5m / README.md
Compactbot's picture
Add tokenizer_config.json, model.py, .gitattributes and README (renamed copy of gpt-s2.5-5m)
5f887f5 verified
|
Raw
History Blame
2.49 kB
---
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
- tiny
- tiny-lm
- tiny-model
- slm
- small-language-model
- from-scratch
- gpt
- gqa
- swiglu
- rope
- rmsnorm
- cpu-trained
library_name: transformers
metrics:
- accuracy
model-index:
- name: HellaSwag
type: text-generation
results: []
---
# compacttest-5m
**5.11M-parameter subword language model, trained from scratch on CPU.**
> This repo is a **renamed copy** of [`Compactbot/gpt-s2.5-5m`](https://huggingface.co/Compactbot/gpt-s2.5-5m),
> renamed at the request of @Datdanboi25 (see `Compactbot/model-requests` #2).
> Weights, config and architecture are **identical** to the original.
## What it is
A small but genuine from-scratch causal LM in the GPT-X2.5 style:
- **Architecture:** GQA (8 query / 2 key-value heads) + SwiGLU FFN + RoPE + RMSNorm,
pre-norm, no biases, **weight-tied** input/output embeddings (the embedding matrix
doubles as the LM head — no separate `lm_head` tensor).
- **Params:** 5,114,112 (verified against the checkpoint).
- **Tokenizer:** 8,192-vocab byte-level BPE (custom, not a standard HF tokenizer).
- **Trained:** on CPU, ~94M TinyStories tokens, cosine LR schedule with warmup.
## How to load
This is a self-contained `nn.Module`, not a transformers-native architecture.
```python
import torch
from model import Model
m = Model()
from safetensors.torch import load_file
sd = load_file("model.safetensors")
m.load_state_dict(sd)
m.eval()
# greedy generation
with torch.no_grad():
x = torch.tensor([[1]])
for _ in range(60):
logits = m(x[:, -512:])[:, -1, :]
x = torch.cat([x, logits.argmax(-1, keepdim=True)], dim=1)
```
## Honest scope
- **Greedy-coherent, sampling-fragile.** At 5M params the model produces readable
prose under greedy decoding but degrades noticeably under sampling. That is the
expected behaviour at this scale, not a bug.
- **Intelligence index:** 0.032 (see the original card for the full eval breakdown).
- It is a reference build demonstrating that a 5M-param from-scratch subword LM is
trainable and coherent on CPU. It is not a chat model.
## Files
| File | Purpose |
|---|---|
| `config.json` | Architecture config (matches `model.py` exactly) |
| `model.py` | Self-contained architecture definition |
| `model.safetensors` | Weights (5,114,112 params, F32) |
| `tokenizer.json` | 8192-vocab BPE tokenizer |
| `tokenizer_config.json` | Tokenizer metadata |