AI for Media Collection NVIDIA AI for Media is a collection that enhance audio, video, and augmented reality effects for media and entertainment workflows • 2 items • Updated 3 days ago • 6
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published 14 days ago • 18
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published 18 days ago • 49
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published 14 days ago • 18
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published 14 days ago • 18
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published 18 days ago • 49
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published about 1 month ago • 4
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published about 1 month ago • 4
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published about 1 month ago • 4
One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications Paper • 2606.25621 • Published Jun 24 • 20
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Paper • 2606.18216 • Published Jun 16 • 65
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 113
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 113
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 113
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding Paper • 2605.19846 • Published May 20 • 3