diffrecon-rl-iter34
Co-evolve iter-1 RL checkpoint iter_34 (Qwen3.5-35B-A3B).
- Run: iter-1 RL (GRPO + dynamic sampling), harness =
diff-reconcile, init = vanilla Qwen3.5-35B-A3B - This is checkpoint 34 of 50 from that run.
Held-out SWE-bench Verified (/500, temp/top_p 0.95, 3-pass)
| harness | score | n |
|---|---|---|
| diff-reconcile | 67.9 +/- 1.5 | 5 |
| combo_fb | 68.1 +/- 1.7 | 3 |
For reference, iter_49 from the same run: 68.2 +/- 0.8 (diff-reconcile, n=6), 68.8 +/- 1.2 (combo_fb, n=10 pooled). The checkpoints are statistically tied on SWE-V.
In-training eval at this step
swe_val 65.3% (128/196), swe_multi 63.4% (single-pass proxy sets).
Not to be confused with
sweagent/rl-combo-iter34 -- that is combo2's iter_34, a different run.
This repo is the iter-1 diff-reconcile RL run.
- Downloads last month
- 7
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support