jsbai-aaron commited on
Commit
10efdfc
·
verified ·
1 Parent(s): e8d0ab4

Add HarmBench safety row: 96.9% refusal (159 standard behaviors, official classifier, hard-harm categories clean)

Browse files
Files changed (1) hide show
  1. README.md +4 -1
README.md CHANGED
@@ -40,14 +40,17 @@ We reserved 121 real software bugs that the model never saw during training. Bef
40
  | **Instruction-following** (IFEval) | 84.66 | 87.21 | **86.37** |
41
  | MMLU-Pro | 64.0% | 70.0% | **66.85%** |
42
  | Terminal-Bench 1.0 (core, 80 tasks) | 33.8% | 33.8% | **18.8%** |
 
43
 
44
  The instruction-following score *improved* over the base model. The coding gains cost nothing on general quality. Gains of this kind usually trade one for the other.
45
 
 
 
46
  The NVFP4 quantization preserves instruction-following (within ~1 point of BF16) and trades real coding capability: Live-60 drops from 21.7% to 15.0%. The BF16 remains the best model; this variant trades that margin for a 40% smaller footprint and ~5GB VRAM.
47
 
48
  *12/32 on a 32-instance subset (the same slice our comparisons use). The quantization costs roughly half the generalization capability of the BF16.
49
 
50
- Note: Terminal-Bench 2.1, safety verification (HarmBench refusal testing), and future benchmark additions are evaluated on the BF16 release (Aztec-Coder-4B). Community quants are welcome; see the BF16 model card for full benchmark coverage.
51
 
52
  Our decontamination protocol is published with the model: none of these benchmark problems overlap the training data.
53
 
 
40
  | **Instruction-following** (IFEval) | 84.66 | 87.21 | **86.37** |
41
  | MMLU-Pro | 64.0% | 70.0% | **66.85%** |
42
  | Terminal-Bench 1.0 (core, 80 tasks) | 33.8% | 33.8% | **18.8%** |
43
+ | HarmBench (harmful-behavior refusal rate, 159 standard behaviors) | 99.4% | 98.1% | **96.9%** |
44
 
45
  The instruction-following score *improved* over the base model. The coding gains cost nothing on general quality. Gains of this kind usually trade one for the other.
46
 
47
+ Safety alignment survived both the RL training and the quantization: the refusal rate on HarmBench's 159 standard harmful behaviors holds at 96.9% (vs 98.1% for the BF16 and 99.4% for the base model), with the serious harm categories clean on all three. See the BF16 model card for the full safety verification.
48
+
49
  The NVFP4 quantization preserves instruction-following (within ~1 point of BF16) and trades real coding capability: Live-60 drops from 21.7% to 15.0%. The BF16 remains the best model; this variant trades that margin for a 40% smaller footprint and ~5GB VRAM.
50
 
51
  *12/32 on a 32-instance subset (the same slice our comparisons use). The quantization costs roughly half the generalization capability of the BF16.
52
 
53
+ Note: Terminal-Bench 2.1 and future benchmark additions are evaluated on the BF16 release (Aztec-Coder-4B). Safety verification (HarmBench refusal testing) is evaluated on both releases. Community quants are welcome; see the BF16 model card for full benchmark coverage.
54
 
55
  Our decontamination protocol is published with the model: none of these benchmark problems overlap the training data.
56