AutoArabic: A Three-Stage Framework for Localizing Video-Text Retrieval Benchmarks Paper • 2509.16438 • Published Sep 19, 2025 • 1
Generative Late-Interaction Embeddings For Visual Document Retrieval Paper • 2609.11808 • Published 5 days ago • 26
GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs Paper • 2605.10762 • Published May 11 • 4
view article Article You could have designed state of the art positional encoding FL33TW00D-HF • Nov 25, 2024 • 500
VideoAtlas: Navigating Long-Form Video in Logarithmic Compute Paper • 2603.17948 • Published Mar 18 • 3
Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers Paper • 2511.01617 • Published Nov 3, 2025 • 3