SalArt-VQA: Diagnosing Whether VLMs Understand Salient Artifacts in Generated Images Paper • 2606.12671 • Published Jun 10 • 1
WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild Paper • 2605.01018 • Published May 1 • 9
Aria: An Open Multimodal Native Mixture-of-Experts Model Paper • 2410.05993 • Published Oct 8, 2024 • 111