jsbai-aaron commited on
Commit
20eaa79
·
verified ·
1 Parent(s): 5052d07

Correct generalization row: audit found 27 of 121 gate instances had entered the training pool pre-measurement; honest split (94 never-trained: 68.1%/60.8%; 27 overlapping: 81.5%/73.6%) replaces the inflated 82.9%. Correction note added with methodology disclosure.

Browse files
Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -33,7 +33,7 @@ We reserved 121 real software bugs that the model never saw during training. Bef
33
 
34
  | Benchmark | Qwen3.5-4B (base) | **Aztec-Coder-4B** | Aztec-Coder-4B-NVFP4 |
35
  |---|---|---|---|
36
- | **Generalization test** (121 unseen bugs, tests run to verify) | 10.1% | **82.9%** | 38.0%* |
37
  | **Live-60** (60 real-world engineering tasks, solved end-to-end in containers) | 15.0% | **21.7%** | 15.0% |
38
  | **Instruction-following** (IFEval) | 84.66 | **87.21** | 86.37 |
39
  | MMLU-Pro | 64.0% | **70.0%** | 66.85% |
@@ -66,7 +66,7 @@ Three stages, each with a plain-language summary:
66
 
67
  2. **Reinforcement learning on real bugs.** The model then practiced on 237 curated software engineering problems: ones it could sometimes solve, but not reliably. For each problem, the model repeatedly attempted a fix. Solutions that made the real hidden tests pass were reinforced; failures were not. This phase, 145 batches of on-policy GRPO, built the actual problem-solving ability.
68
 
69
- 3. **Generalization checks.** At every stage boundary, we re-tested the model on problems it had never trained on. The 10.1% to 82.9% jump above is the result.
70
 
71
  ### Training data
72
 
 
33
 
34
  | Benchmark | Qwen3.5-4B (base) | **Aztec-Coder-4B** | Aztec-Coder-4B-NVFP4 |
35
  |---|---|---|---|
36
+ | **Generalization test** (121 never-trained bugs, tests run to verify) | 10.1% | **68.1%** (instance) / 60.8% (per-attempt) on the 94 fully held-out; 81.5%/73.6% on 27 later-found overlapping | 38.0%* |
37
  | **Live-60** (60 real-world engineering tasks, solved end-to-end in containers) | 15.0% | **21.7%** | 15.0% |
38
  | **Instruction-following** (IFEval) | 84.66 | **87.21** | 86.37 |
39
  | MMLU-Pro | 64.0% | **70.0%** | 66.85% |
 
66
 
67
  2. **Reinforcement learning on real bugs.** The model then practiced on 237 curated software engineering problems: ones it could sometimes solve, but not reliably. For each problem, the model repeatedly attempted a fix. Solutions that made the real hidden tests pass were reinforced; failures were not. This phase, 145 batches of on-policy GRPO, built the actual problem-solving ability.
68
 
69
+ 3. **Generalization checks.** At every stage boundary, we re-tested the model on problems it had never trained on. *Correction (2026-09-29): an earlier version of this card reported 82.9% on the full 121-instance set. A post-release audit found 27 of those instances had entered the training pool before the measurement, inflating the number. The corrected figures above separate the 94 genuinely never-trained instances (68.1% instance-level, 60.8% per-attempt) from the 27 overlapping ones (81.5%/73.6%). The training gains remain large: the same 94 instances went from 7.0% before training to 60.8% after.*
70
 
71
  ### Training data
72