anurag051194 commited on
Commit
7fbd4b1
·
verified ·
1 Parent(s): 221b22c

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +2 -1
README.md CHANGED
@@ -167,7 +167,8 @@ truncation artefact rather than a measured score; excluding the row, its 13-data
167
 
168
  Lengths marked ≈ are derived from stored generations using each model's characters-per-token
169
  ratio rather than re-tokenized directly; the method reproduces the three directly measured
170
- lengths to within 3.5%.
 
171
 
172
  The honest summary of this table is that TwIL-LM3 does not lead it. Larger models score higher,
173
  in order of size, and the 120B leads nine of fourteen rows. Two things are worth extracting
 
167
 
168
  Lengths marked ≈ are derived from stored generations using each model's characters-per-token
169
  ratio rather than re-tokenized directly; the method reproduces the three directly measured
170
+ lengths to within 3.5%. All the models are evaluated on our evaluation harness with 300 randomly shuffled samples (consistent for all the models) from each of the datasets.
171
+ MuSR used all 756 samples. The results might vary on different test set sizes but are statistically significant for comparisons.
172
 
173
  The honest summary of this table is that TwIL-LM3 does not lead it. Larger models score higher,
174
  in order of size, and the 120B leads nine of fourteen rows. Two things are worth extracting