Text Generation
Transformers
Safetensors
English
gpt-s2.5
tiny
tiny-lm
tiny-model
slm
small-language-model
from-scratch
gpt
gqa
swiglu
rope
rmsnorm
cpu-trained
Instructions to use Compactbot/compacttest-5m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Compactbot/compacttest-5m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Compactbot/compacttest-5m")# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Compactbot/compacttest-5m", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Compactbot/compacttest-5m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Compactbot/compacttest-5m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Compactbot/compacttest-5m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Compactbot/compacttest-5m
- SGLang
How to use Compactbot/compacttest-5m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Compactbot/compacttest-5m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Compactbot/compacttest-5m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Compactbot/compacttest-5m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Compactbot/compacttest-5m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Compactbot/compacttest-5m with Docker Model Runner:
docker model run hf.co/Compactbot/compacttest-5m
Add tokenizer_config.json, model.py, .gitattributes and README (renamed copy of gpt-s2.5-5m)
5f887f5 verified metadata
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
- tiny
- tiny-lm
- tiny-model
- slm
- small-language-model
- from-scratch
- gpt
- gqa
- swiglu
- rope
- rmsnorm
- cpu-trained
library_name: transformers
metrics:
- accuracy
model-index:
- name: HellaSwag
type: text-generation
results: []
compacttest-5m
5.11M-parameter subword language model, trained from scratch on CPU.
This repo is a renamed copy of
Compactbot/gpt-s2.5-5m, renamed at the request of @Datdanboi25 (seeCompactbot/model-requests#2). Weights, config and architecture are identical to the original.
What it is
A small but genuine from-scratch causal LM in the GPT-X2.5 style:
- Architecture: GQA (8 query / 2 key-value heads) + SwiGLU FFN + RoPE + RMSNorm,
pre-norm, no biases, weight-tied input/output embeddings (the embedding matrix
doubles as the LM head — no separate
lm_headtensor). - Params: 5,114,112 (verified against the checkpoint).
- Tokenizer: 8,192-vocab byte-level BPE (custom, not a standard HF tokenizer).
- Trained: on CPU, ~94M TinyStories tokens, cosine LR schedule with warmup.
How to load
This is a self-contained nn.Module, not a transformers-native architecture.
import torch
from model import Model
m = Model()
from safetensors.torch import load_file
sd = load_file("model.safetensors")
m.load_state_dict(sd)
m.eval()
# greedy generation
with torch.no_grad():
x = torch.tensor([[1]])
for _ in range(60):
logits = m(x[:, -512:])[:, -1, :]
x = torch.cat([x, logits.argmax(-1, keepdim=True)], dim=1)
Honest scope
- Greedy-coherent, sampling-fragile. At 5M params the model produces readable prose under greedy decoding but degrades noticeably under sampling. That is the expected behaviour at this scale, not a bug.
- Intelligence index: 0.032 (see the original card for the full eval breakdown).
- It is a reference build demonstrating that a 5M-param from-scratch subword LM is trainable and coherent on CPU. It is not a chat model.
Files
| File | Purpose |
|---|---|
config.json |
Architecture config (matches model.py exactly) |
model.py |
Self-contained architecture definition |
model.safetensors |
Weights (5,114,112 params, F32) |
tokenizer.json |
8192-vocab BPE tokenizer |
tokenizer_config.json |
Tokenizer metadata |