E6E831728 commited on
Commit
3163a06
·
verified ·
1 Parent(s): baf891e

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +106 -0
README.md CHANGED
@@ -80,6 +80,112 @@ with torch.no_grad():
80
  print(tokenizer.decode(output_ids[0].tolist()))
81
  ```
82
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
83
  ## Intended use
84
 
85
  This checkpoint is provided for anonymous review and reproducibility of the paper's main claim: a trainable input embedding table is not necessary for useful language modeling in the studied regime.
 
80
  print(tokenizer.decode(output_ids[0].tolist()))
81
  ```
82
 
83
+ ## Standardized base-model evaluation
84
+
85
+ The checkpoint was evaluated as a base causal language model with
86
+ EleutherAI LM Evaluation Harness `v0.4.10`.
87
+
88
+ Evaluation protocol:
89
+
90
+ - Hugging Face backend: `hf`
91
+ - maximum context length: 1,024
92
+ - `add_bos_token=False`
93
+ - no chat template
94
+ - deterministic likelihood-based evaluation
95
+ - harness seeds: `0,1234,1234,1234`
96
+ - base checkpoints only; no SFT or instruction checkpoints
97
+
98
+ | Metric | Learned input table | Fixed Binary-16 | Affine GF(2), table-free | SmolLM2-135M | SmolLM2-360M |
99
+ |---|---:|---:|---:|---:|---:|
100
+ | HellaSwag acc | 28.49 ± 0.45 | 29.04 ± 0.45 | 29.04 ± 0.45 | 35.36 ± 0.48 | 43.05 ± 0.49 |
101
+ | HellaSwag acc_norm | 31.32 ± 0.46 | 32.32 ± 0.47 | 31.80 ± 0.46 | 43.02 ± 0.49 | 56.28 ± 0.50 |
102
+ | ARC-Easy acc | 46.38 ± 1.02 | 47.90 ± 1.03 | 47.64 ± 1.02 | 64.44 ± 0.98 | 70.24 ± 0.94 |
103
+ | ARC-Easy acc_norm | 40.70 ± 1.01 | 40.87 ± 1.01 | 41.20 ± 1.01 | 58.75 ± 1.01 | 68.18 ± 0.96 |
104
+ | ARC-Challenge acc | 20.39 ± 1.18 | 19.62 ± 1.16 | 21.33 ± 1.20 | 28.07 ± 1.31 | 36.26 ± 1.40 |
105
+ | ARC-Challenge acc_norm | 25.85 ± 1.28 | 26.19 ± 1.28 | 24.83 ± 1.26 | 29.61 ± 1.33 | 38.05 ± 1.42 |
106
+ | PIQA acc | 62.35 ± 1.13 | 62.57 ± 1.13 | 62.68 ± 1.13 | 68.44 ± 1.08 | 71.38 ± 1.05 |
107
+ | PIQA acc_norm | 60.61 ± 1.14 | 62.08 ± 1.13 | 60.94 ± 1.14 | 68.39 ± 1.08 | 71.82 ± 1.05 |
108
+ | WinoGrande acc | 50.20 ± 1.41 | 50.12 ± 1.41 | 50.43 ± 1.41 | 52.57 ± 1.40 | 59.35 ± 1.38 |
109
+ | OpenBookQA acc | 18.40 ± 1.73 | 17.20 ± 1.69 | 17.60 ± 1.70 | 22.00 ± 1.85 | 24.80 ± 1.93 |
110
+ | OpenBookQA acc_norm | 29.20 ± 2.04 | 31.00 ± 2.07 | 29.40 ± 2.04 | 32.60 ± 2.10 | 37.80 ± 2.17 |
111
+ | CommonsenseQA acc | 20.31 ± 1.15 | 19.90 ± 1.14 | 20.23 ± 1.15 | 19.90 ± 1.14 | 21.05 ± 1.17 |
112
+ | MMLU 0-shot | 24.13 ± 0.36 | 23.86 ± 0.36 | 24.11 ± 0.36 | 24.24 ± 0.36 | 25.47 ± 0.37 |
113
+ | MMLU 5-shot | 25.68 ± 0.37 | 25.60 ± 0.37 | 25.66 ± 0.37 | 25.39 ± 0.37 | 25.05 ± 0.37 |
114
+ | LAMBADA accuracy | 22.38 ± 0.58 | 21.23 ± 0.57 | 21.99 ± 0.58 | 42.97 ± 0.69 | 53.31 ± 0.70 |
115
+ | LAMBADA perplexity | 95.14 ± 4.01 | 101.74 ± 4.27 | 100.61 ± 4.17 | 19.06 ± 0.63 | 9.38 ± 0.27 |
116
+ | WikiText word perplexity | 81.04 | 74.87 | 76.17 | 25.53 | 18.84 |
117
+ | WikiText byte perplexity | 2.27 | 2.24 | 2.25 | 1.83 | 1.73 |
118
+ | WikiText bits/byte | 1.19 | 1.16 | 1.17 | 0.87 | 0.79 |
119
+
120
+ The three paper checkpoints form the controlled architectural comparison.
121
+ SmolLM2-135M and SmolLM2-360M are external reference models, not matched
122
+ baselines: they use different architectures, tokenizers, training mixtures,
123
+ and much larger pretraining budgets. SmolLM2-135M was trained on approximately
124
+ 2T tokens and SmolLM2-360M on approximately 4T tokens, whereas the paper
125
+ checkpoints saw approximately 16–17B tokens. Their scores therefore provide
126
+ context for absolute capability and must not be interpreted as isolating the
127
+ effect of the input parameterization.
128
+
129
+ Perplexity values should be interpreted especially cautiously across different
130
+ tokenizers. The primary controlled comparison is among the three paper models,
131
+ which share the same tokenizer, data pipeline, and architecture.
132
+
133
+
134
+ ## Input-interface audit
135
+
136
+ This checkpoint stores a deterministic `65,536 × 16` binary codebook as a
137
+ frozen `nn.Embedding` for computational convenience. During training, the
138
+ table was initialized from the fixed token codes, marked with
139
+ `requires_grad=False`, and excluded from the optimizer. It therefore contained
140
+ 1,048,576 stored but non-trainable values and contributed zero trainable input
141
+ parameters.
142
+
143
+ The released checkpoint can be audited directly:
144
+
145
+ ```python
146
+ import torch
147
+ from transformers import AutoModelForCausalLM
148
+
149
+ repo_id = "E6E831728/fixed-minimal-binary-code"
150
+
151
+ model = AutoModelForCausalLM.from_pretrained(
152
+ repo_id,
153
+ trust_remote_code=True,
154
+ torch_dtype=torch.float32,
155
+ ).cpu().eval()
156
+
157
+ embedding = model.get_input_embeddings()
158
+ weight = embedding.weight.detach()
159
+
160
+ vocab_size, code_bits = weight.shape
161
+ ids = torch.arange(vocab_size, dtype=torch.long)
162
+ positions = torch.arange(code_bits, dtype=torch.long)
163
+
164
+ expected = ((ids[:, None] >> positions[None, :]) & 1).float()
165
+ expected[model.config.pad_token_id].zero_()
166
+
167
+ print("shape:", tuple(weight.shape))
168
+ print("unique values:", torch.unique(weight).tolist())
169
+ print("all entries binary:", bool(torch.all((weight == 0) | (weight == 1))))
170
+ print("exact canonical-code match:", bool(torch.equal(weight, expected)))
171
+ print("mismatching entries:", int((weight != expected).sum().item()))
172
+
173
+ assert tuple(weight.shape) == (65536, 16)
174
+ assert torch.all((weight == 0) | (weight == 1))
175
+ assert torch.equal(weight, expected)
176
+ ```
177
+
178
+ Expected audit properties:
179
+
180
+ ```text
181
+ shape: (65536, 16)
182
+ unique values: [0.0, 1.0]
183
+ all entries binary: True
184
+ exact canonical-code match: True
185
+ mismatching entries: 0
186
+ ```
187
+
188
+
189
  ## Intended use
190
 
191
  This checkpoint is provided for anonymous review and reproducibility of the paper's main claim: a trainable input embedding table is not necessary for useful language modeling in the studied regime.