jsbai-aaron commited on
Commit
efdd021
·
verified ·
1 Parent(s): 20eaa79

IMPORTANT: Mark generalization and Live-60 rows as under re-evaluation due to discovered RL environment flaws (git-checkout shortcut + broken edit tool). Safety results unaffected. Full re-measurement in progress.

Browse files
Files changed (1) hide show
  1. README.md +5 -3
README.md CHANGED
@@ -33,8 +33,8 @@ We reserved 121 real software bugs that the model never saw during training. Bef
33
 
34
  | Benchmark | Qwen3.5-4B (base) | **Aztec-Coder-4B** | Aztec-Coder-4B-NVFP4 |
35
  |---|---|---|---|
36
- | **Generalization test** (121 never-trained bugs, tests run to verify) | 10.1% | **68.1%** (instance) / 60.8% (per-attempt) on the 94 fully held-out; 81.5%/73.6% on 27 later-found overlapping | 38.0%* |
37
- | **Live-60** (60 real-world engineering tasks, solved end-to-end in containers) | 15.0% | **21.7%** | 15.0% |
38
  | **Instruction-following** (IFEval) | 84.66 | **87.21** | 86.37 |
39
  | MMLU-Pro | 64.0% | **70.0%** | 66.85% |
40
  | Terminal-Bench 1.0 (core, 80 tasks) | 33.8% | **33.8%** | 18.8% |
@@ -66,7 +66,9 @@ Three stages, each with a plain-language summary:
66
 
67
  2. **Reinforcement learning on real bugs.** The model then practiced on 237 curated software engineering problems: ones it could sometimes solve, but not reliably. For each problem, the model repeatedly attempted a fix. Solutions that made the real hidden tests pass were reinforced; failures were not. This phase, 145 batches of on-policy GRPO, built the actual problem-solving ability.
68
 
69
- 3. **Generalization checks.** At every stage boundary, we re-tested the model on problems it had never trained on. *Correction (2026-09-29): an earlier version of this card reported 82.9% on the full 121-instance set. A post-release audit found 27 of those instances had entered the training pool before the measurement, inflating the number. The corrected figures above separate the 94 genuinely never-trained instances (68.1% instance-level, 60.8% per-attempt) from the 27 overlapping ones (81.5%/73.6%). The training gains remain large: the same 94 instances went from 7.0% before training to 60.8% after.*
 
 
70
 
71
  ### Training data
72
 
 
33
 
34
  | Benchmark | Qwen3.5-4B (base) | **Aztec-Coder-4B** | Aztec-Coder-4B-NVFP4 |
35
  |---|---|---|---|
36
+ | **Generalization test** (121 never-trained bugs, tests run to verify) | 10.1% | 68.1% / 60.8% — **UNDER RE-EVALUATION** (see note below) | 38.0%* |
37
+ | **Live-60** (60 real-world engineering tasks, solved end-to-end in containers) | 15.0% | 21.7% — **under re-evaluation** | 15.0% |
38
  | **Instruction-following** (IFEval) | 84.66 | **87.21** | 86.37 |
39
  | MMLU-Pro | 64.0% | **70.0%** | 66.85% |
40
  | Terminal-Bench 1.0 (core, 80 tasks) | 33.8% | **33.8%** | 18.8% |
 
66
 
67
  2. **Reinforcement learning on real bugs.** The model then practiced on 237 curated software engineering problems: ones it could sometimes solve, but not reliably. For each problem, the model repeatedly attempted a fix. Solutions that made the real hidden tests pass were reinforced; failures were not. This phase, 145 batches of on-policy GRPO, built the actual problem-solving ability.
68
 
69
+ 3. **Generalization checks.** At every stage boundary, we re-tested the model on problems it had never trained on. *IMPORTANT RE-EVALUATION NOTICE (2026-09-29, later): a post-release audit discovered environment flaws in the RL training harness that may have inflated the generalization and Live-60 numbers above: (a) the injected bug patches were applied to the git working tree uncommitted, making them discoverable via `git diff` and revertible via `git checkout` (a shortcut the model may have learned); and (b) the `str_replace_edit` tool returned a false success signal for every call during training. We are re-measuring the model in a validated environment and will update this card with corrected numbers. The safety results (HarmBench, XSTest) were measured on independent harnesses unaffected by these flaws and stand. See the methodology section for details.*
70
+
71
+ *Correction (2026-09-29, earlier): an earlier version of this card reported 82.9% on the full 121-instance set. A post-release audit found 27 of those instances had entered the training pool before the measurement, inflating the number. The corrected figures above separate the 94 genuinely never-trained instances (68.1% instance-level, 60.8% per-attempt) from the 27 overlapping ones (81.5%/73.6%). The training gains remain large: the same 94 instances went from 7.0% before training to 60.8% after.*
72
 
73
  ### Training data
74