anurag051194 commited on
Commit
588354c
·
verified ·
1 Parent(s): a8247e4

VibeThinker-webAI-trained: add estimated tok/s

Browse files
Files changed (1) hide show
  1. README.md +8 -3
README.md CHANGED
@@ -129,7 +129,7 @@ Cells marked † need the engine note below.
129
  | **macro gate** | **0.5539** | 0.4313 | 0.4118 | 0.508 | 0.4466 | 0.4218 | 0.2925 | 0.3473 | 0.3757 | 0.5336 | — |
130
  | **strict-7** | **0.2879** | 0.1821 | 0.2021 | — | 0.1121 | 0.1971 | 0.1229 | 0.1579 | 0.1714 | 0.2093 | — |
131
  | macro_primary | **0.5875** | 0.4825 | 0.4637 | 0.579 | 0.4313 | 0.4475 | 0.3450 | 0.4188 | 0.4213 | 0.5750 | — |
132
- | tok/s | 21169 † | 21916 † | 28010 † | — | 11097 † | 15880 | 16160 | 25230 | 22480 | 9420 | 3374 |
133
  | mean gen length | 1902 | 2951 | 2688 | — | 2486 | **564** | 696 | 2296 | 1830 | 2094 | 1005 |
134
  | **ans/s** | 11.1 † | 7.4 † | 10.4 † | — | 4.5 † | **28.1** | 23.2 | 10.9 | 12.0 | 4.5 | 3.4 |
135
 
@@ -153,7 +153,12 @@ family tables score `mcq_answer` with loose-match credit (0.695), so the strict
153
  six-lane average and strict-7 — which all need strict scoring — are left as —. Its perplexities
154
  come from that separate run; the family table records the untuned VibeThinker-3B at 18.70 and
155
  21.80 there, against 16.93 and 27.03 in the columns above, so read them within their own source.
156
- Throughput and generation length were not measured for it.
 
 
 
 
 
157
 
158
  † **Throughput for TwIL-LM3-Pro, its base and VibeThinker-3B** was measured with the same
159
  dedicated protocol and prompt file as the peer columns (128 prompts × 512 generated tokens, EOS
@@ -241,7 +246,7 @@ spots in absolute terms are `procedural` (strict 0.1200, loose 0.2350) and FOL t
241
  | math500 | 0.7467 | 0.6567 | 0.7900 | 0.790 | 0.3600 ¶¶¶ | 0.6900 | 0.4233 | 0.7133 | 0.7800 | 0.6100 | **0.8433** |
242
  | **macro (10 CoT datasets)** | 0.7901 | 0.7942 | 0.8097 | 0.802 | 0.7683 | 0.7339 | 0.6997 | 0.7523 | 0.7884 | 0.8493 | **0.8689** |
243
  | **macro (all 14)** | 0.7425 | 0.7332 | 0.7262 | 0.728 | 0.6611 | 0.6694 | 0.6245 | 0.6814 | 0.7378 | 0.7591 | **0.8086** |
244
- | tok/s | 21169 † | 21916 † | 28010 † | — | 11097 † | 15880 | 16160 | 25230 | 22480 | 9420 | 3374 |
245
  | mean gen length | ≈792 | ≈1282 | ≈1789 | — | ≈2787 | **482** | 510 | ≈796 | ≈1327 | ≈1931 | 801 |
246
  | **ans/s** | 26.7 † | 17.1 † | ≈15.7 † | — | ≈4.0 † | **32.9** | 31.7 | ≈31.7 | ≈16.9 | 4.9 | 4.2 |
247
 
 
129
  | **macro gate** | **0.5539** | 0.4313 | 0.4118 | 0.508 | 0.4466 | 0.4218 | 0.2925 | 0.3473 | 0.3757 | 0.5336 | — |
130
  | **strict-7** | **0.2879** | 0.1821 | 0.2021 | — | 0.1121 | 0.1971 | 0.1229 | 0.1579 | 0.1714 | 0.2093 | — |
131
  | macro_primary | **0.5875** | 0.4825 | 0.4637 | 0.579 | 0.4313 | 0.4475 | 0.3450 | 0.4188 | 0.4213 | 0.5750 | — |
132
+ | tok/s | 21169 † | 21916 † | 28010 † | ≈28010 ★ † | 11097 † | 15880 | 16160 | 25230 | 22480 | 9420 | 3374 |
133
  | mean gen length | 1902 | 2951 | 2688 | — | 2486 | **564** | 696 | 2296 | 1830 | 2094 | 1005 |
134
  | **ans/s** | 11.1 † | 7.4 † | 10.4 † | — | 4.5 † | **28.1** | 23.2 | 10.9 | 12.0 | 4.5 | 3.4 |
135
 
 
153
  six-lane average and strict-7 — which all need strict scoring — are left as —. Its perplexities
154
  come from that separate run; the family table records the untuned VibeThinker-3B at 18.70 and
155
  21.80 there, against 16.93 and 27.03 in the columns above, so read them within their own source.
156
+ Its `tok/s` is an **estimate, not a measurement**: merging changes weights but not architecture,
157
+ parameter count or tokenizer, and the throughput protocol fixes the output at 512 generated tokens
158
+ with EOS ignored, so the decode rate does not depend on what the weights say. It is therefore set
159
+ equal to the 28,010 tok/s measured for the base VibeThinker-3B in the same session (a ≈ and both
160
+ marks in the cell). Mean generation length, and so `ans/s`, does depend on the weights and was not
161
+ measured, so those rows stay —.
162
 
163
  † **Throughput for TwIL-LM3-Pro, its base and VibeThinker-3B** was measured with the same
164
  dedicated protocol and prompt file as the peer columns (128 prompts × 512 generated tokens, EOS
 
246
  | math500 | 0.7467 | 0.6567 | 0.7900 | 0.790 | 0.3600 ¶¶¶ | 0.6900 | 0.4233 | 0.7133 | 0.7800 | 0.6100 | **0.8433** |
247
  | **macro (10 CoT datasets)** | 0.7901 | 0.7942 | 0.8097 | 0.802 | 0.7683 | 0.7339 | 0.6997 | 0.7523 | 0.7884 | 0.8493 | **0.8689** |
248
  | **macro (all 14)** | 0.7425 | 0.7332 | 0.7262 | 0.728 | 0.6611 | 0.6694 | 0.6245 | 0.6814 | 0.7378 | 0.7591 | **0.8086** |
249
+ | tok/s | 21169 † | 21916 † | 28010 † | ≈28010 ★ † | 11097 † | 15880 | 16160 | 25230 | 22480 | 9420 | 3374 |
250
  | mean gen length | ≈792 | ≈1282 | ≈1789 | — | ≈2787 | **482** | 510 | ≈796 | ≈1327 | ≈1931 | 801 |
251
  | **ans/s** | 26.7 † | 17.1 † | ≈15.7 † | — | ≈4.0 † | **32.9** | 31.7 | ≈31.7 | ≈16.9 | 4.9 | 4.2 |
252