SlowGuess/ABForge-Qwen3-8B
Text Generation • 8B • Updated • 380 • 1
Paper-grounded ablation design. Unified ABForge-Qwen3-8B (both tasks, one checkpoint), its ablations, and the data.
Note Released unified model: one checkpoint for both tasks (SFT -> GRPO, update 200). AblationBench 55.9 / 62.4.
Note Unified SFT-only stage (mixed 1:1, 1 full epoch / step 602) - the RL initialization. 30.7 / 52.2.
Note Unified RL-only ablation: GRPO directly from Qwen3-8B, no SFT warm start. 52.2 / 54.9.
Note SFT + RL training pools, the AblationBench eval sets, and per-paper judge outputs for all 21 evaluated models.