When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 1 day ago • 30
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 4 days ago • 76
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement Paper • 2609.13406 • Published 7 days ago • 60
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work Paper • 2609.11977 • Published 14 days ago • 159
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 7 days ago • 311
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 4 days ago • 299
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 8 days ago • 207
DataFlex-RL: An Evaluation Platform for RLVR Data Policies Paper • 2609.06107 • Published 13 days ago • 125
eslab1234/smolvla_multitask_5blocks_v3_684ep_from575_285k_vlmexpert_20k Robotics • 0.5B • Updated 5 days ago • 28 • 1
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 10 days ago • 422
Dr. Claw: An AI Scientist Workspace for Vibe Research Paper • 2609.00365 • Published 18 days ago • 183
RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning Paper • 2609.03199 • Published 16 days ago • 122
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 17 days ago • 269