LH-Tech-AI commited on
Commit
3efcd43
·
verified ·
1 Parent(s): 83bc631

Update we-were-wrong-GDN-acc-norm.html

Browse files
Files changed (1) hide show
  1. we-were-wrong-GDN-acc-norm.html +36 -0
we-were-wrong-GDN-acc-norm.html CHANGED
@@ -173,7 +173,43 @@
173
  </div>
174
 
175
  <div class="post-content">
 
 
176
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
177
  </div>
178
  </article>
179
 
 
173
  </div>
174
 
175
  <div class="post-content">
176
+ <p><strong>Yes.</strong> We were wrong. And we are honest about it. Read this to see us explaining what went wrong.</p>
177
+ <p>So, two days ago, we wrote this blog showing almost unbelievable performance on our new GatedDeltaNet 5M model. This model actually wasn't as good as we thought. It was much worse.</p>
178
 
179
+ <h2>The acc vs acc_norm trap</h2>
180
+ <p>As always, we were evaluating the model against our typical LM-Eval tasks (PIQA, HellaSwag, ARC-Easy and ARC-Challenge) and it went pretty straight-forward.<br>
181
+ But we accidently forgot to implement using acc_<strong>norm</strong> into the benchmark script and blindly took acc (with out normalization 😭).</p>
182
+
183
+ <h2>The new model and the correct results</h2>
184
+ <p>We do not have the weights for the ~300M tokens model anymore. Sorry. But we retrained and improved the model - now on 5B high-quality webdata tokens. Here are the <strong>REAL</strong> results, this time with acc_norm!</p>
185
+
186
+ <div class="table-wrap">
187
+ <table>
188
+ <thead>
189
+ <tr><th>Model</th><th>Train tokens</th><th>ARC-Easy</th><th>ARC-Challenge</th><th>HellaSwag</th><th>PIQA</th></tr>
190
+ </thead>
191
+ <tbody>
192
+ <tr><td><strong>Supra-5M-GatedDeltaNet (CURRENT ONE)</strong></td><td><strong>5B</strong></td><td><strong>33.59%</strong></td><td><strong>23.21%</strong></td><td><strong>26.74%</strong></td><td><strong>52.88%</strong></td></tr>
193
+ <tr><td>fromziro/Qana-mini-5M</td><td>~21B</td><td>34.97%</td><td>23.21%</td><td>27.60%</td><td>57.18%</td></tr>
194
+ <tr><td>AxiomicLabs/GPT-S2-5M</td><td>~75B</td><td>33.92%</td><td>22.87%</td><td>27.87%</td><td>57.56%</td></tr>
195
+ <tr><td>User01110/CMA-8M</td><td>~21B</td><td>35.35%</td><td>23.29%</td><td>28.19%</td><td>58.22%</td></tr>
196
+ </tbody>
197
+ </table>
198
+ </div>
199
+
200
+ <p>That's the whole truth. Now with acc_norm.</p>
201
+
202
+ <h2>The new model</h2>
203
+ <p>As mentioned above, we have trained the model entirely new from scratch - this time on 5B tokens.</p>
204
+
205
+ <div class="callout">
206
+ <span>Link to the model: <a style="color: var(--accent) !important;" href="https://huggingface.co/SupraLabs/SupraGDN-5M">https://huggingface.co/SupraLabs/SupraGDN-5M</a></span>
207
+ </div>
208
+
209
+ <h2>What this means for the soon upcoming Supra3</h2>
210
+ <p>This architecture changes a lot and we'll take a deeper look at it to maybe even implement it into our future work - maybe even into Supra3!</p>
211
+ <p><strong>Stay tuned!</strong></p>
212
+
213
  </div>
214
  </article>
215