EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 2 days ago • 227
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Paper • 2608.13417 • Published 9 days ago • 52
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published 27 days ago • 106
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play Paper • 2509.25541 • Published Sep 29, 2025 • 142
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning Paper • 2509.22576 • Published Sep 26, 2025 • 137