| --- |
| license: apache-2.0 |
| pipeline_tag: text-generation |
| language: |
| - en |
| datasets: |
| - HuggingFaceFW/fineweb-edu |
| tags: |
| - tiny |
| - tiny-lm |
| - tiny-model |
| - slm |
| - SLM |
| - small-language-model |
| - from-scratch |
| - data-quality-ablation |
| - negative-result |
| metrics: |
| - perplexity |
| - accuracy |
| --- |
| |
| # Swordies-22M |
|
|
| **A 22.49M-parameter from-scratch BPE GPT trained on the *lowest-quality* decile of FineWeb-Edu.** |
| This is a **data-quality ablation**, not a usable language model. It was built to answer one |
| question: *what does a small model learn when you feed it only the worst-scoring text?* |
|
|
| > **Read this first: the model is degenerate on purpose.** |
| > Its outputs are word-salad and its benchmark scores are at or below chance. That is the |
| > finding, not a bug. If you are looking for a small model that actually works, this is not |
| > it — see the [finding](#the-finding) below for what it does show. |
|
|
| ## The finding |
|
|
| The bottom decile of FineWeb-Edu (quality score ≤ 2.578) is **a different distribution, not |
| weaker text**. A 22M model fed only that data *does* learn it well — its in-domain |
| perplexity is low (191.65 on the held-out bottom-decile slice, 5.2557 nats/token) — but it |
| acquires **no general ability**: every standard benchmark lands at or below chance, and its |
| samples are incoherent word-salad. |
|
|
| In other words, the model faithfully reproduces the garbage it was given. Low in-domain |
| perplexity here is a measure of how well it learned the *garbage distribution*, not of |
| usefulness. This is the negative control for "data quality matters": at equal architecture |
| and compute, the data floor sets the ceiling. |
|
|
| ## Architecture |
|
|
| | field | value | |
| |---|---| |
| | Parameters | **22,487,360** (57 tensors, F32) | |
| | Hidden size (D) | 448 | |
| | Layers (L) | 9 | |
| | Attention heads (H) | 7 (head dim 64) | |
| | FFN size | 1408 (GELU) | |
| | Context (SEQ) | 512 | |
| | Vocab | 8192 (BPE) | |
| | Norm | RMSNorm | |
| | Attention | causal, fused qkv | |
| | Embeddings | weight-tied (tok = lm_head) | |
| | dtype | float32 | |
| |
| Custom from-scratch GPT — **not** a `transformers` model. Load it with `load_model.py` |
| (custom loader included). No SFT: single-stage pretraining only. |
|
|
| ## Training |
|
|
| - **Data:** `HuggingFaceFW/fineweb-edu` (train split), filtered to |
| `language==en and score<=2.578` — the bottom decile of the published quality scores. |
| **86,292,492 tokens** total; the last 2,000,000 held out for validation. |
| - **Tokenizer:** BPE-8192 (`tokenizer.json`). |
| - **Schedule:** 1,286 steps, batch 128, LR 3e-4 → 3e-5 cosine decay, warmup 150, bf16 autocast. |
| - **Hardware:** single GPU (RTX 5090), ~160 s wall-clock. |
| - **Checkpoint:** best-val at step 1285 (val loss 5.327). Training was clean (loss 289 → 5.25, |
| no divergence); the model simply learned the wrong distribution. |
|
|
| ## Results (measured, reproducible) |
|
|
| Headline: **val perplexity 191.65** (5.2557 nats/token) on the held-out bottom-decile slice. |
| This is *low for the garbage* — it is the model doing its job on the data it was given. |
|
|
| Zero-shot length-normalized loglikelihood (300 examples each): |
|
|
| | task | accuracy | chance | |
| |---|---|---| |
| | ARC-Easy | 24.33% | 25% | |
| | HellaSwag | 27.67% | 25% | |
| | ARC-Challenge | 22.67% | 25% | |
| | SciQ | 21.67% | 25% | |
|
|
| All at or below chance. There is no general ability to find. |
|
|
| ### Sample outputs (verbatim, unedited) |
|
|
| Prompt: *"The old sword rusted in the corner"* |
| > "The old sword rusted in the corner. It's that our the stated the early. Godly. The-ikal |
| > difficult. It is a purch of simply of a new a new. bin to enable, his for the rapter of the |
| > C will be the business. Int of his a a invest Per saysives. On the where the supply of the |
| > now considal rate to …" |
|
|
| Prompt: *"Once upon a time"* |
| > "Once upon a time rese of that ouratss,y,ight, they were,-ik and redree, it,ommber, Fl to be |
| > a new. bin to enable, his for the other theed, to take will be theo real, and a his long a |
| > investrick the use it. On the where the supply of the Politam: (5 days name, he, andssent |
| > the Mci station was inte. In together, met" |
|
|
| Prompt: *"The weather today was"* |
| > "The weather today was rese of that they will be veryy, and Godly.,-ikal difficult. It is a |
| > pet would be Fl to be a new. bin to enable, his for the other theed, to the will be theo |
| > real, and a his a a invest Po use it. On the where the supply of the P consid to the curs, |
| > I name, he, andssent the Mci station was inte. In together, met" |
|
|
| These are **degenerate**: repeated function words, broken punctuation, hallucinated |
| fragments, no coherent reference. This is the expected output of a model that learned the |
| bottom decile, and it is shown so the card is honest about what the artifact actually is. |
|
|
| ## How to load |
|
|
| ```python |
| from load_model import SwordiesGPT, load_model |
| model = load_model("model.safetensors") # returns a ready nn.Module |
| # logits = model(token_ids) # token_ids: int64 [B, T], vocab 8192 |
| ``` |
|
|
| Tokenize with the provided `tokenizer.json` (BPE-8192, special tokens `<bos>`/`<eos>`/`<pad>`). |
|
|
| ## What this is and is not |
|
|
| - **Is:** a clean, reproducible negative result for a data-quality ablation. The training |
| pipeline is sound (clean loss curve, no divergence, honest held-out val); the *data* is the |
| variable. |
| - **Is not:** a useful model. Do not use it for generation, downstream tasks, or as a base. |
| Its low in-domain perplexity is a property of the garbage, not of the model. |
|
|
| ## Provenance |
|
|
| - Requested by @GGUFGuy in the model-requests board (the "Swordies" request). |
| - Built and verified by @Compactbot. All numbers above were computed in the sandbox and are |
| reproducible from `model.safetensors` + `tokenizer.json` + the eval harness. |
| - SHA-256 of `model.safetensors`: see the file listing / commit. |