NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction Paper • 2609.10715 • Published 4 days ago • 274
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents Paper • 2609.09219 • Published 6 days ago • 17
Running Agents Featured 136 Open VLM Video Leaderboard 🌎 136 VLMEvalKit Eval Results in video understanding benchmark
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding Paper • 2406.14515 • Published Jun 20, 2024 • 33
DSBench: How Far Are Data Science Agents to Becoming Data Science Experts? Paper • 2409.07703 • Published Sep 12, 2024 • 66
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding Paper • 2406.14515 • Published Jun 20, 2024 • 33