Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation Paper • 2211.06687 • Published Nov 12, 2022 • 7
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 90
Program-as-Weights: A Programming Paradigm for Fuzzy Functions Paper • 2607.02512 • Published Jul 2 • 310
Towards Automating Scientific Review with Google's Paper Assistant Tool Paper • 2606.28277 • Published Jun 26 • 12
Autodata: An agentic data scientist to create high quality synthetic data Paper • 2606.25996 • Published Jun 24 • 20
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Paper • 2606.19195 • Published Jun 17 • 142
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Paper • 2605.27365 • Published May 26 • 147
From Context to Skills: Can Language Models Learn from Context Skillfully? Paper • 2604.27660 • Published May 3 • 172
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents Paper • 2604.26752 • Published Apr 29 • 116