--- license: apache-2.0 pipeline_tag: text-generation language: en tags: - tiny - tiny-lm - tiny-model - slm - small-language-model - from-scratch - gpt - gqa - swiglu - rope - rmsnorm - cpu-trained library_name: transformers metrics: - accuracy model-index: - name: HellaSwag type: text-generation results: [] --- # compacttest-5m **5.11M-parameter subword language model, trained from scratch on CPU.** > This repo is a **renamed copy** of [`Compactbot/gpt-s2.5-5m`](https://huggingface.co/Compactbot/gpt-s2.5-5m), > renamed at the request of @Datdanboi25 (see `Compactbot/model-requests` #2). > Weights, config and architecture are **identical** to the original. ## What it is A small but genuine from-scratch causal LM in the GPT-X2.5 style: - **Architecture:** GQA (8 query / 2 key-value heads) + SwiGLU FFN + RoPE + RMSNorm, pre-norm, no biases, **weight-tied** input/output embeddings (the embedding matrix doubles as the LM head — no separate `lm_head` tensor). - **Params:** 5,114,112 (verified against the checkpoint). - **Tokenizer:** 8,192-vocab byte-level BPE (custom, not a standard HF tokenizer). - **Trained:** on CPU, ~94M TinyStories tokens, cosine LR schedule with warmup. ## How to load This is a self-contained `nn.Module`, not a transformers-native architecture. ```python import torch from model import Model m = Model() from safetensors.torch import load_file sd = load_file("model.safetensors") m.load_state_dict(sd) m.eval() # greedy generation with torch.no_grad(): x = torch.tensor([[1]]) for _ in range(60): logits = m(x[:, -512:])[:, -1, :] x = torch.cat([x, logits.argmax(-1, keepdim=True)], dim=1) ``` ## Honest scope - **Greedy-coherent, sampling-fragile.** At 5M params the model produces readable prose under greedy decoding but degrades noticeably under sampling. That is the expected behaviour at this scale, not a bug. - **Intelligence index:** 0.032 (see the original card for the full eval breakdown). - It is a reference build demonstrating that a 5M-param from-scratch subword LM is trainable and coherent on CPU. It is not a chat model. ## Files | File | Purpose | |---|---| | `config.json` | Architecture config (matches `model.py` exactly) | | `model.py` | Self-contained architecture definition | | `model.safetensors` | Weights (5,114,112 params, F32) | | `tokenizer.json` | 8192-vocab BPE tokenizer | | `tokenizer_config.json` | Tokenizer metadata |