anurag051194 commited on
Commit
5d90f3a
·
verified ·
1 Parent(s): 7fbd4b1

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +3 -2
README.md CHANGED
@@ -167,8 +167,9 @@ truncation artefact rather than a measured score; excluding the row, its 13-data
167
 
168
  Lengths marked ≈ are derived from stored generations using each model's characters-per-token
169
  ratio rather than re-tokenized directly; the method reproduces the three directly measured
170
- lengths to within 3.5%. All the models are evaluated on our evaluation harness with 300 randomly shuffled samples (consistent for all the models) from each of the datasets.
171
- MuSR used all 756 samples. The results might vary on different test set sizes but are statistically significant for comparisons.
 
172
 
173
  The honest summary of this table is that TwIL-LM3 does not lead it. Larger models score higher,
174
  in order of size, and the 120B leads nine of fourteen rows. Two things are worth extracting
 
167
 
168
  Lengths marked ≈ are derived from stored generations using each model's characters-per-token
169
  ratio rather than re-tokenized directly; the method reproduces the three directly measured
170
+ lengths to within 3.5%. All the models are evaluated on our evaluation harness with 300 randomly shuffled
171
+ samples (samples are consistent and same for all the models) from each of the datasets for quick compute. The results might
172
+ vary on different test set sizes but the comparitive accuracies are statistically significant across models.
173
 
174
  The honest summary of this table is that TwIL-LM3 does not lead it. Larger models score higher,
175
  in order of size, and the 120B leads nine of fourteen rows. Two things are worth extracting