Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails Paper • 2609.09134 • Published 13 days ago • 11
Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure Paper • 2608.08722 • Published Aug 9 • 8
view article Article Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL +6 aminediroHF, qgallouedec, kashif, lewtun, edbeeching, albertvillanova, lvwerra, sergiopaniego • May 27 • 50
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 21 days ago • 40
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO Paper • 2608.27351 • Published 25 days ago • 22
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Paper • 2608.20061 • Published Aug 20 • 47
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Paper • 2608.16590 • Published Aug 17 • 150
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Paper • 2608.17253 • Published Aug 19 • 97
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published Aug 18 • 108
Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution Paper • 2608.07645 • Published Aug 7 • 27
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 95
view article Article Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident +2 hlarcher, XciD, raphael-gl, chris-rannou • Jul 27 • 504
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published Jul 22 • 33
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published Jul 23 • 107