Curated papers, models, datasets, and demos for AI-agent runtime safety, prompt injection, MCP security, and tool-call guardrails.
-
MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits
Paper • 2504.03767 • Published • 3 -
Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems
Paper • 2601.17548 • Published -
ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
Paper • 2601.10156 • Published • 26 -
Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use
Paper • 2603.03205 • Published • 13