Omega-S: A Functional Resilience Index for LLM Fine-Tuning Paper • 2608.03887 • Published 9 days ago • 6
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design Paper • 2608.10299 • Published 3 days ago • 122
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Paper • 2608.11079 • Published 2 days ago • 12
MatrAIx: Simulating the World with 8.3 Billion Persona Agents Paper • 2608.04205 • Published 9 days ago • 34
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Paper • 2608.03573 • Published 7 days ago • 51
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution Paper • 2608.08311 • Published 5 days ago • 79
Characterizing the Quality Profile of AI-Generated C++ in Production Paper • 2608.06640 • Published 7 days ago • 10
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Paper • 2608.00782 • Published 12 days ago • 16
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Paper • 2608.03764 • Published 9 days ago • 27
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Paper • 2608.05139 • Published 8 days ago • 27
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents Paper • 2608.05013 • Published 9 days ago • 35
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published 7 days ago • 59
ChronoVision: Temporal Reasoning via Latent State Reconstruction Paper • 2608.05631 • Published 7 days ago • 39
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Paper • 2608.06301 • Published 7 days ago • 34
Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval Paper • 2608.06060 • Published 7 days ago • 40
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Paper • 2608.00155 • Published 13 days ago • 25
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Paper • 2608.02499 • Published 10 days ago • 24