jsbai-aaron commited on
Commit
f7efec7
·
verified ·
1 Parent(s): 10efdfc

Add XSTest row: 99.8% benign compliance — the full three-variant safety table

Browse files
Files changed (1) hide show
  1. README.md +2 -1
README.md CHANGED
@@ -41,10 +41,11 @@ We reserved 121 real software bugs that the model never saw during training. Bef
41
  | MMLU-Pro | 64.0% | 70.0% | **66.85%** |
42
  | Terminal-Bench 1.0 (core, 80 tasks) | 33.8% | 33.8% | **18.8%** |
43
  | HarmBench (harmful-behavior refusal rate, 159 standard behaviors) | 99.4% | 98.1% | **96.9%** |
 
44
 
45
  The instruction-following score *improved* over the base model. The coding gains cost nothing on general quality. Gains of this kind usually trade one for the other.
46
 
47
- Safety alignment survived both the RL training and the quantization: the refusal rate on HarmBench's 159 standard harmful behaviors holds at 96.9% (vs 98.1% for the BF16 and 99.4% for the base model), with the serious harm categories clean on all three. See the BF16 model card for the full safety verification.
48
 
49
  The NVFP4 quantization preserves instruction-following (within ~1 point of BF16) and trades real coding capability: Live-60 drops from 21.7% to 15.0%. The BF16 remains the best model; this variant trades that margin for a 40% smaller footprint and ~5GB VRAM.
50
 
 
41
  | MMLU-Pro | 64.0% | 70.0% | **66.85%** |
42
  | Terminal-Bench 1.0 (core, 80 tasks) | 33.8% | 33.8% | **18.8%** |
43
  | HarmBench (harmful-behavior refusal rate, 159 standard behaviors) | 99.4% | 98.1% | **96.9%** |
44
+ | XSTest (benign-but-scary request compliance, 450 prompts) | 54.0% | 97.1% | **99.8%** |
45
 
46
  The instruction-following score *improved* over the base model. The coding gains cost nothing on general quality. Gains of this kind usually trade one for the other.
47
 
48
+ Safety alignment survived both the RL training and the quantization: the refusal rate on HarmBench's 159 standard harmful behaviors holds at 96.9% (vs 98.1% for the BF16 and 99.4% for the base model), with the serious harm categories clean on all three. On the over-refusal side the quant answers 99.8% of benign-but-scary requests (vs 54.0% for the base model) — the deployment artifact's best safety number. See the BF16 model card for the full safety verification. See the BF16 model card for the full safety verification.
49
 
50
  The NVFP4 quantization preserves instruction-following (within ~1 point of BF16) and trades real coding capability: Live-60 drops from 21.7% to 15.0%. The BF16 remains the best model; this variant trades that margin for a 40% smaller footprint and ~5GB VRAM.
51