LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 4 days ago • 176
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published 16 days ago • 39
TACO: Tool-Augmented Credit Optimization for Agentic Tool Use Paper • 2606.30251 • Published 21 days ago • 22
ActiveMimic: Egocentric Video Pretraining with Active Perception Paper • 2606.06194 • Published Jun 4 • 2
Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization Paper • 2605.28109 • Published May 27 • 23
ankile/real01c-insert-marker-d1-baseline-uniform-r1-25k-nocf-eval-sobol20 Viewer • Updated May 28 • 5.22k • 32 • 1
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 171