MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use Paper • 2512.24565 • Published Jan 21
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Paper • 2604.16706 • Published Apr 17 • 1
GuardianAgentBench: Where Agents Fail and How to Guard Them Paper • 2607.20982 • Published Jul 23 • 1