MM-Matryoshka: Towards Budget-Elastic Visual Document Retrieval via a 2D Multimodal Matryoshka Training Framework Paper • 2606.07654 • Published Jun 3
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval Paper • 2604.10167 • Published Apr 11
EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next Paper • 2603.12147 • Published Mar 12
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models Paper • 2502.11051 • Published May 27, 2025
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations Paper • 2603.01666 • Published Mar 2 • 1
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval Paper • 2602.19961 • Published Feb 23 • 2
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework Paper • 2602.19549 • Published Feb 23 • 1
view article Article NEO-unify: Building Native Multimodal Unified Models End to End sensenova • Mar 5 • 179
CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding Paper • 2601.21262 • Published Jan 29 • 1
Aligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Models Paper • 2512.04981 • Published Dec 4, 2025 • 9
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey Paper • 2412.02104 • Published Dec 3, 2024