All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation Paper • 2609.27901 • Published 11 days ago • 22
Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World Paper • 2609.23038 • Published 15 days ago • 63
DeltaWAM: Delta World Action Models for Bimanual Manipulation Paper • 2609.28811 • Published 11 days ago • 19
Coding Agents for Generalized Task and Motion Planning Problems Paper • 2609.30233 • Published 10 days ago • 25
InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data Paper • 2609.31394 • Published 9 days ago • 29
EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks Paper • 2609.28236 • Published 11 days ago • 39
LEGO-Anything: Coding Agents for 3D Scene Reconstruction Paper • 2609.36380 • Published 6 days ago • 130
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation Paper • 2610.02201 • Published 3 days ago • 26
World Observer: Joint Actor-Observer Generation for Persistent World Modeling Paper • 2610.02162 • Published 3 days ago • 71
Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding Paper • 2609.32019 • Published 9 days ago • 40