IVGT: Implicit Visual Geometry Transformer for Neural Scene Representation Paper • 2605.16258 • Published May 21
R2RDreamer: 3D-aware Data Augmentation for Spatially-generalized 2D Manipulation Policies Paper • 2606.17040 • Published Jun 15
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 2 days ago • 109
Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment Paper • 2605.08064 • Published May 8 • 1
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 2 days ago • 109
iMaC: Translating Actions into Motion and Contact Images for Embodied World Models Paper • 2606.09813 • Published Jun 8 • 13
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 433
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents Paper • 2604.07430 • Published Apr 8 • 181
DVD: Deterministic Video Depth Estimation with Generative Priors Paper • 2603.12250 • Published Mar 12 • 28
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training Paper • 2603.12255 • Published Mar 12 • 91
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion Paper • 2507.02813 • Published Jul 3, 2025 • 60
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Paper • 2505.23747 • Published May 29, 2025 • 69
Running on Zero Agents Featured 262 MatchAnything 🏢 262 Find similar images and match them across collections
VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step Paper • 2504.01956 • Published Apr 2, 2025 • 41
The GAN is dead; long live the GAN! A Modern GAN Baseline Paper • 2501.05441 • Published Jan 9, 2025 • 99
DreamCinema: Cinematic Transfer with Free Camera and 3D Character Paper • 2408.12601 • Published Aug 22, 2024 • 32
ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model Paper • 2408.16767 • Published Aug 29, 2024 • 32