Update README.md
Browse files
README.md
CHANGED
|
@@ -25,8 +25,6 @@ Built on Qwen3-0.6B. Trained with **mechanical supervision** (NVIDIA Aegis-2.0 l
|
|
| 25 |
| Aegis-2.0 test F1 (in-domain detection) | **76%** (P72/R81) | NVIDIA's primary in-domain metric; in the range of 8B guards |
|
| 26 |
| ToxicChat F1 (out-of-domain) | 28% | OOD generalization; see limitations |
|
| 27 |
|
| 28 |
-
Incumbent sub-1B guards score ~0.7% on policy-flip. Railz-R adds interpretable reasoning at no incumbent's-size cost.
|
| 29 |
-
|
| 30 |
## Prompt format (verdict-first, then reasoning)
|
| 31 |
|
| 32 |
```
|
|
|
|
| 25 |
| Aegis-2.0 test F1 (in-domain detection) | **76%** (P72/R81) | NVIDIA's primary in-domain metric; in the range of 8B guards |
|
| 26 |
| ToxicChat F1 (out-of-domain) | 28% | OOD generalization; see limitations |
|
| 27 |
|
|
|
|
|
|
|
| 28 |
## Prompt format (verdict-first, then reasoning)
|
| 29 |
|
| 30 |
```
|