Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers Paper • 2610.00531 • Published 10 days ago • 60
Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers Paper • 2610.00531 • Published 10 days ago • 60
EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery Paper • 2609.40340 • Published 10 days ago • 111
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions Paper • 2608.30428 • Published Aug 31 • 14
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions Paper • 2608.30428 • Published Aug 31 • 14
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions Paper • 2608.30428 • Published Aug 31 • 14
MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models Paper • 2510.16641 • Published Oct 18, 2025 • 5
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games Paper • 2509.01052 • Published Sep 1, 2025 • 22
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games Paper • 2509.01052 • Published Sep 1, 2025 • 22
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games Paper • 2509.01052 • Published Sep 1, 2025 • 22 • 1
ChartCap: Mitigating Hallucination of Dense Chart Captioning Paper • 2508.03164 • Published Aug 5, 2025 • 7
ChartCap: Mitigating Hallucination of Dense Chart Captioning Paper • 2508.03164 • Published Aug 5, 2025 • 7
ChartCap: Mitigating Hallucination of Dense Chart Captioning Paper • 2508.03164 • Published Aug 5, 2025 • 7 • 2
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games Paper • 2506.03610 • Published Jun 4, 2025 • 10