File size: 2,485 Bytes
5f887f5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
---
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
  - tiny
  - tiny-lm
  - tiny-model
  - slm
  - small-language-model
  - from-scratch
  - gpt
  - gqa
  - swiglu
  - rope
  - rmsnorm
  - cpu-trained
library_name: transformers
metrics:
  - accuracy
model-index:
  - name: HellaSwag
    type: text-generation
    results: []
---

# compacttest-5m

**5.11M-parameter subword language model, trained from scratch on CPU.**

> This repo is a **renamed copy** of [`Compactbot/gpt-s2.5-5m`](https://huggingface.co/Compactbot/gpt-s2.5-5m),
> renamed at the request of @Datdanboi25 (see `Compactbot/model-requests` #2).
> Weights, config and architecture are **identical** to the original.

## What it is

A small but genuine from-scratch causal LM in the GPT-X2.5 style:

- **Architecture:** GQA (8 query / 2 key-value heads) + SwiGLU FFN + RoPE + RMSNorm,
  pre-norm, no biases, **weight-tied** input/output embeddings (the embedding matrix
  doubles as the LM head — no separate `lm_head` tensor).
- **Params:** 5,114,112 (verified against the checkpoint).
- **Tokenizer:** 8,192-vocab byte-level BPE (custom, not a standard HF tokenizer).
- **Trained:** on CPU, ~94M TinyStories tokens, cosine LR schedule with warmup.

## How to load

This is a self-contained `nn.Module`, not a transformers-native architecture.

```python
import torch
from model import Model

m = Model()
from safetensors.torch import load_file
sd = load_file("model.safetensors")
m.load_state_dict(sd)
m.eval()

# greedy generation
with torch.no_grad():
    x = torch.tensor([[1]])
    for _ in range(60):
        logits = m(x[:, -512:])[:, -1, :]
        x = torch.cat([x, logits.argmax(-1, keepdim=True)], dim=1)
```

## Honest scope

- **Greedy-coherent, sampling-fragile.** At 5M params the model produces readable
  prose under greedy decoding but degrades noticeably under sampling. That is the
  expected behaviour at this scale, not a bug.
- **Intelligence index:** 0.032 (see the original card for the full eval breakdown).
- It is a reference build demonstrating that a 5M-param from-scratch subword LM is
  trainable and coherent on CPU. It is not a chat model.

## Files

| File | Purpose |
|---|---|
| `config.json` | Architecture config (matches `model.py` exactly) |
| `model.py` | Self-contained architecture definition |
| `model.safetensors` | Weights (5,114,112 params, F32) |
| `tokenizer.json` | 8192-vocab BPE tokenizer |
| `tokenizer_config.json` | Tokenizer metadata |