Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation Paper • 2608.10932 • Published 13 days ago
WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models Paper • 2608.04964 • Published 19 days ago • 13
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models Paper • 2605.25077 • Published May 24 • 22
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning Paper • 2605.21988 • Published May 21 • 1
A Molecular Multimodal Foundation Model Associating Molecule Graphs with Natural Language Paper • 2209.05481 • Published Sep 12, 2022
GUI Agents with Reinforcement Learning: Toward Digital Inhabitants Paper • 2604.27955 • Published Apr 30 • 2
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion Paper • 2603.06140 • Published Mar 6 • 1