What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents Paper • 2610.06406 • Published 6 days ago • 20
AdSpark: A Large-Scale Dataset and Benchmark for Product-Centric Advertisement Video Generation Paper • 2610.10047 • Published 4 days ago • 17
World Editing: Intervening on Executable Worlds at Increasing Depth Paper • 2610.02331 • Published 10 days ago • 30
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF Image-Text-to-Text • 177B • Updated 12 days ago • 4.19M • 785
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model Paper • 2609.18323 • Published 25 days ago • 120
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published Sep 7 • 376
Beyond Solver Verdicts: Generative Reward Models for Autoformalization Paper • 2609.11085 • Published Sep 10 • 35
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation Paper • 2609.11486 • Published Sep 10 • 34