Update README.md
Browse files
README.md
CHANGED
|
@@ -52,6 +52,10 @@ Official evaluations were conducted using the **EleutherAI LM Evaluation Harness
|
|
| 52 |
| **GSM8K (Flexible Extract)** | Exact Match (Regex Clean) | **32.45%** | Mathematical Thought & Resolution |
|
| 53 |
| **GSM8K (Strict)** | Exact Match (Rigid Parse) | **19.79%** | Formatted Mathematical Output |
|
| 54 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 55 |
### 🔍 Comparative Engineering Insights
|
| 56 |
|
| 57 |
* **Punching Above Weight Classes:** Atomight-V2.1-0.5B outpaces Meta's larger **Llama-3.2-1B-Instruct** on localized logic-retrieval metrics, clearing **59.3%** on ARC-Easy and **33.8%** on ARC-Challenge compared to Llama's *56.7%* and *31.8%* respectively.
|
|
|
|
| 52 |
| **GSM8K (Flexible Extract)** | Exact Match (Regex Clean) | **32.45%** | Mathematical Thought & Resolution |
|
| 53 |
| **GSM8K (Strict)** | Exact Match (Rigid Parse) | **19.79%** | Formatted Mathematical Output |
|
| 54 |
|
| 55 |
+
<p align="center">
|
| 56 |
+
<img src="OfficialBenchmarkAtomight2.1.png" alt="Atomight V2.1 Benchmark" width="500" style="max-width: 100%;">
|
| 57 |
+
</p>
|
| 58 |
+
|
| 59 |
### 🔍 Comparative Engineering Insights
|
| 60 |
|
| 61 |
* **Punching Above Weight Classes:** Atomight-V2.1-0.5B outpaces Meta's larger **Llama-3.2-1B-Instruct** on localized logic-retrieval metrics, clearing **59.3%** on ARC-Easy and **33.8%** on ARC-Challenge compared to Llama's *56.7%* and *31.8%* respectively.
|