StephYang/qwen3.5-4b-swe-opd-tmax500-baseline-step105 Text Generation • 4B • Updated 10 days ago • 405
StephYang/qwen3.5-4b-swe-opd-tmax500-baseline-step105 Text Generation • 4B • Updated 10 days ago • 405
Lego-RL Collection Harness-native RL for coding agents: the trained policy and the training task index. • 9 items • Updated 3 days ago • 11
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents Paper • 2608.17393 • Published about 1 month ago • 25