| --- |
| license: other |
| base_model: Qwen/Qwen3.5-35B-A3B |
| --- |
| |
| # diffrecon-rl-iter34 |
|
|
| Co-evolve **iter-1 RL** checkpoint **iter_34** (Qwen3.5-35B-A3B). |
| |
| - Run: iter-1 RL (GRPO + dynamic sampling), harness = `diff-reconcile`, init = vanilla Qwen3.5-35B-A3B |
| - This is checkpoint **34 of 50** from that run. |
| |
| ## Held-out SWE-bench Verified (/500, temp/top_p 0.95, 3-pass) |
| |
| | harness | score | n | |
| |---|---|---| |
| | diff-reconcile | 67.9 +/- 1.5 | 5 | |
| | combo_fb | 68.1 +/- 1.7 | 3 | |
| |
| For reference, iter_49 from the same run: 68.2 +/- 0.8 (diff-reconcile, n=6), |
| 68.8 +/- 1.2 (combo_fb, n=10 pooled). The checkpoints are statistically tied on SWE-V. |
| |
| ## In-training eval at this step |
| |
| swe_val 65.3% (128/196), swe_multi 63.4% (single-pass proxy sets). |
| |
| ## Not to be confused with |
| |
| `sweagent/rl-combo-iter34` -- that is **combo2's** iter_34, a different run. |
| This repo is the **iter-1 diff-reconcile** RL run. |
|
|