Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 11 days ago • 689
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published Aug 15 • 447
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Paper • 2608.03979 • Published Aug 4 • 54
MMSkills: Towards Multimodal Skills for General Visual Agents Paper • 2605.13527 • Published May 14 • 123
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards Paper • 2605.10899 • Published May 11 • 78
MolmoAct2: Action Reasoning Models for Real-world Deployment Paper • 2605.02881 • Published May 4 • 356