jsbai-aaron commited on
Commit
f6ca4bd
·
verified ·
1 Parent(s): f6a24c3

Add NVFP4 HarmBench safety cell: 96.9% refusal (quant noise pushes 2 borderline cases; hard-harm categories clean)

Browse files
Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -39,11 +39,11 @@ We reserved 121 real software bugs that the model never saw during training. Bef
39
  | MMLU-Pro | 64.0% | **70.0%** | 66.85% |
40
  | Terminal-Bench 1.0 (core, 80 tasks) | 33.8% | **33.8%** | 18.8% |
41
  | Terminal-Bench 2.1 (89 tasks, both models, same protocol) | 14.6% | 11.2% | not evaluated |
42
- | HarmBench (harmful-behavior refusal rate, 159 standard behaviors) | 99.4% | **98.1%** | not evaluated |
43
 
44
  The instruction-following score *improved* over the base model. The coding gains cost nothing on general quality. Gains of this kind usually trade one for the other.
45
 
46
- We also verified that safety alignment survived training. Each judge-confirmed refusal test used HarmBench's standard set of 159 harmful behaviors. The base model refuses 99.4% of them; our model refuses 98.1%. The serious harm categories (chemical and biological, illegal activity, harassment) are clean on both models. The few requests each model does answer are edge cases, like writing a persuasive article about a disputed topic, and they barely overlap between the two models. The refusals were confirmed by two independent runs of the official HarmBench classifier, with identical results.
47
 
48
  *NVFP4 generalization: 12/32 on a 32-instance subset (the same slice our comparisons use). The quantization costs roughly half the generalization capability.
49
 
 
39
  | MMLU-Pro | 64.0% | **70.0%** | 66.85% |
40
  | Terminal-Bench 1.0 (core, 80 tasks) | 33.8% | **33.8%** | 18.8% |
41
  | Terminal-Bench 2.1 (89 tasks, both models, same protocol) | 14.6% | 11.2% | not evaluated |
42
+ | HarmBench (harmful-behavior refusal rate, 159 standard behaviors) | 99.4% | **98.1%** | 96.9% |
43
 
44
  The instruction-following score *improved* over the base model. The coding gains cost nothing on general quality. Gains of this kind usually trade one for the other.
45
 
46
+ We also verified that safety alignment survived training. Each judge-confirmed refusal test used HarmBench's standard set of 159 harmful behaviors. The base model refuses 99.4% of them; our model refuses 98.1%, and the NVFP4 quantized version refuses 96.9%. The serious harm categories (chemical and biological, illegal activity, harassment) are clean on all three models. The few requests each model does answer are edge cases, like writing a persuasive article about a disputed topic, and they barely overlap between the models. The refusals were confirmed by two independent runs of the official HarmBench classifier, with identical results.
47
 
48
  *NVFP4 generalization: 12/32 on a 32-instance subset (the same slice our comparisons use). The quantization costs roughly half the generalization capability.
49