MBZUAI/Omni-Embed-Mini-0.9B
Feature Extraction • Updated • 35 • 1
Natural Language Processing, Machine Learning, and Computer Vision
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
Training-Free Speech-Centric Omni Understanding with Frozen VLMs