Mage Collection A family of lightweight multimodal models, including understanding and generation. • 8 items • Updated 13 days ago • 24
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 13 days ago • 36
nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Viewer • Updated Jun 9 • 156M • 36.7k • 37
view article Article Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action nvidia • Jun 1 • 87
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Paper • 2506.05414 • Published Jun 4, 2025 • 4