Update README.md
Browse files
README.md
CHANGED
|
@@ -95,9 +95,9 @@ Curated (not scraped) from frontier 2025-era open datasets across STEM, science,
|
|
| 95 |
|
| 96 |
---
|
| 97 |
|
| 98 |
-
##
|
| 99 |
|
| 100 |
-
These results
|
| 101 |
|
| 102 |
| Benchmark | Metric | Score |
|
| 103 |
|---|---:|---:|
|
|
@@ -114,7 +114,7 @@ These results are self-reported for `NovatasticRoScript/Atomight-V2.5-1.7B`, eva
|
|
| 114 |
|
| 115 |
---
|
| 116 |
|
| 117 |
-
## How it compares with other small language models (we recommend verifying it as the other data from other models came from a third-party sources)
|
| 118 |
|
| 119 |
Scores below for other models are drawn from their respective model cards / technical reports, not re-run by us. Provided for context only — evaluation harnesses and prompt formats differ across labs, so treat this as directional rather than exact.
|
| 120 |
|
|
|
|
| 95 |
|
| 96 |
---
|
| 97 |
|
| 98 |
+
## Benchmark results
|
| 99 |
|
| 100 |
+
These are the results of the benchmarks for *Atomight-V2.5-1.7B*, evaluated with `lm-evaluation-harness`.
|
| 101 |
|
| 102 |
| Benchmark | Metric | Score |
|
| 103 |
|---|---:|---:|
|
|
|
|
| 114 |
|
| 115 |
---
|
| 116 |
|
| 117 |
+
## How it compares with other small language models (we recommend verifying it, as the other data from other models came from a third-party sources)
|
| 118 |
|
| 119 |
Scores below for other models are drawn from their respective model cards / technical reports, not re-run by us. Provided for context only — evaluation harnesses and prompt formats differ across labs, so treat this as directional rather than exact.
|
| 120 |
|