Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
validation 's Collections
RAG Validation, Grounding and Data Quality
AI Validation Methods, Benchmarks and Testing
Agent Validation, Tool Use and Autonomous Systems
AI Model Validation, Robustness and Reliability
AI Validation — Models, Agents & Reliability

Agent Validation, Tool Use and Autonomous Systems

updated 8 days ago

Curated resources for validating agent behavior, tool use, trajectories, recovery and autonomous AI systems.

Upvote
-

  • Running

    Agent Validation

    🧭

    Validate agent behavior, tool use, and recovery.


  • MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use

    Paper • 2512.24565 • Published Jan 21

  • Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench

    Paper • 2604.16706 • Published Apr 17 • 1

  • An Empirical Study of Automating Agent Evaluation

    Paper • 2605.11378 • Published May 12 • 2

  • GuardianAgentBench: Where Agents Fail and How to Guard Them

    Paper • 2607.20982 • Published Jul 23 • 1

  • aiagentkarl/agent-evaluation-benchmark

    Viewer • Updated Mar 20 • 57 • 26

  • fahadhafeezofficial/agent-failure-bench

    Viewer • Updated 15 days ago • 12k • 412

  • validation/ai-validation-checklists

    Viewer • Updated 8 days ago • 60 • 68
Upvote
-
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs