Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement Paper • 2609.13406 • Published 7 days ago • 60
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement Paper • 2609.13406 • Published 7 days ago • 60
Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem Paper • 2608.11654 • Published Aug 12 • 1
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement Paper • 2609.13406 • Published 7 days ago • 60
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought Paper • 2606.15753 • Published Jul 28
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published 24 days ago • 144
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 171
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Paper • 2606.11324 • Published Jun 9 • 174
Bridging Evolutionary Algorithms and Reinforcement Learning: A Comprehensive Survey on Hybrid Algorithms Paper • 2401.11963 • Published Jan 22, 2024 • 2
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Paper • 2507.06892 • Published Jul 9, 2025 • 1
Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem Paper • 2608.11654 • Published Aug 12 • 1
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Paper • 2507.06892 • Published Jul 9, 2025 • 1
Bridging Evolutionary Algorithms and Reinforcement Learning: A Comprehensive Survey on Hybrid Algorithms Paper • 2401.11963 • Published Jan 22, 2024 • 2
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published 24 days ago • 144
Running Featured 789 Agent Memory Leaderboard 🧠789 Unified memory evaluation · Results expected August 12.
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 171
Trust-Region Behavior Blending for On-Policy Distillation Paper • 2605.31159 • Published May 29 • 69
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Paper • 2606.11324 • Published Jun 9 • 174