Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs Paper • 2609.04753 • Published 5 days ago • 12
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 6 days ago • 288
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 8 days ago • 259
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 7 days ago • 532
Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase Paper • 2608.29310 • Published 11 days ago • 27
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 9 days ago • 144
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs? Paper • 2603.24472 • Published Mar 25 • 59
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning Paper • 2608.14277 • Published 26 days ago • 35
Intern-S2-Preview: Scientific Agentic Foundation Model Paper • 2608.13505 • Published 27 days ago • 72
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published Aug 7 • 111
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper • 2608.05987 • Published Aug 6 • 101
On-Policy Self-Distillation without Any Supervision Paper • 2608.06296 • Published about 1 month ago • 218
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 106