Spaces:
Running
Running
Update we-were-wrong-GDN-acc-norm.html
Browse files
we-were-wrong-GDN-acc-norm.html
CHANGED
|
@@ -173,7 +173,43 @@
|
|
| 173 |
</div>
|
| 174 |
|
| 175 |
<div class="post-content">
|
|
|
|
|
|
|
| 176 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 177 |
</div>
|
| 178 |
</article>
|
| 179 |
|
|
|
|
| 173 |
</div>
|
| 174 |
|
| 175 |
<div class="post-content">
|
| 176 |
+
<p><strong>Yes.</strong> We were wrong. And we are honest about it. Read this to see us explaining what went wrong.</p>
|
| 177 |
+
<p>So, two days ago, we wrote this blog showing almost unbelievable performance on our new GatedDeltaNet 5M model. This model actually wasn't as good as we thought. It was much worse.</p>
|
| 178 |
|
| 179 |
+
<h2>The acc vs acc_norm trap</h2>
|
| 180 |
+
<p>As always, we were evaluating the model against our typical LM-Eval tasks (PIQA, HellaSwag, ARC-Easy and ARC-Challenge) and it went pretty straight-forward.<br>
|
| 181 |
+
But we accidently forgot to implement using acc_<strong>norm</strong> into the benchmark script and blindly took acc (with out normalization 😭).</p>
|
| 182 |
+
|
| 183 |
+
<h2>The new model and the correct results</h2>
|
| 184 |
+
<p>We do not have the weights for the ~300M tokens model anymore. Sorry. But we retrained and improved the model - now on 5B high-quality webdata tokens. Here are the <strong>REAL</strong> results, this time with acc_norm!</p>
|
| 185 |
+
|
| 186 |
+
<div class="table-wrap">
|
| 187 |
+
<table>
|
| 188 |
+
<thead>
|
| 189 |
+
<tr><th>Model</th><th>Train tokens</th><th>ARC-Easy</th><th>ARC-Challenge</th><th>HellaSwag</th><th>PIQA</th></tr>
|
| 190 |
+
</thead>
|
| 191 |
+
<tbody>
|
| 192 |
+
<tr><td><strong>Supra-5M-GatedDeltaNet (CURRENT ONE)</strong></td><td><strong>5B</strong></td><td><strong>33.59%</strong></td><td><strong>23.21%</strong></td><td><strong>26.74%</strong></td><td><strong>52.88%</strong></td></tr>
|
| 193 |
+
<tr><td>fromziro/Qana-mini-5M</td><td>~21B</td><td>34.97%</td><td>23.21%</td><td>27.60%</td><td>57.18%</td></tr>
|
| 194 |
+
<tr><td>AxiomicLabs/GPT-S2-5M</td><td>~75B</td><td>33.92%</td><td>22.87%</td><td>27.87%</td><td>57.56%</td></tr>
|
| 195 |
+
<tr><td>User01110/CMA-8M</td><td>~21B</td><td>35.35%</td><td>23.29%</td><td>28.19%</td><td>58.22%</td></tr>
|
| 196 |
+
</tbody>
|
| 197 |
+
</table>
|
| 198 |
+
</div>
|
| 199 |
+
|
| 200 |
+
<p>That's the whole truth. Now with acc_norm.</p>
|
| 201 |
+
|
| 202 |
+
<h2>The new model</h2>
|
| 203 |
+
<p>As mentioned above, we have trained the model entirely new from scratch - this time on 5B tokens.</p>
|
| 204 |
+
|
| 205 |
+
<div class="callout">
|
| 206 |
+
<span>Link to the model: <a style="color: var(--accent) !important;" href="https://huggingface.co/SupraLabs/SupraGDN-5M">https://huggingface.co/SupraLabs/SupraGDN-5M</a></span>
|
| 207 |
+
</div>
|
| 208 |
+
|
| 209 |
+
<h2>What this means for the soon upcoming Supra3</h2>
|
| 210 |
+
<p>This architecture changes a lot and we'll take a deeper look at it to maybe even implement it into our future work - maybe even into Supra3!</p>
|
| 211 |
+
<p><strong>Stay tuned!</strong></p>
|
| 212 |
+
|
| 213 |
</div>
|
| 214 |
</article>
|
| 215 |
|