--- title: README emoji: 🛡️ colorFrom: indigo colorTo: red sdk: static pinned: false ---
Securing the age of agency.
Website · Hugging Face · Collection · X · LinkedIn · Contact
## About Exponential Security Labs is an AI safety and security research company. We study AI systems that act through tools and multi-step workflows, and build automated red-teaming and adaptive guardrail methods for them. Our team brings nearly a decade of work in adversarial robustness and multimodal AI safety. We publish evaluation datasets and develop methods for identifying failures, measuring harmful assistance, and testing safeguards against adaptive attacks. ## What we work on - **Agent safety evaluation** — benchmarks and evaluation methods for tool-using, multi-step AI systems. - **Automated red-teaming** — agents and methods that identify vulnerabilities in models, tools, and agent workflows. - **Adaptive guardrails** — defences designed to respond to changing attacks and deployment contexts. - **Adversarial robustness** — reliable behaviour under malicious or unexpected inputs across language and vision systems. ## On the Hub Browse the [ExpSec AI Safety Evaluations collection](https://huggingface.co/collections/ExpSec/expsec-ai-safety-evaluations-6a5f7308854e355ca281e04c) for our public evaluation releases. ### [Sovereign-Jbreak](https://huggingface.co/datasets/ExpSec/Sovereign-Jbreak) A regional safety evaluation dataset for testing whether a language model gives actionable help towards a harmful objective. The initial release contains expert-authored evaluation items for Europe and India, with goal-specific rubrics and no model responses or attack trajectories. > **Content warning:** This dataset describes high-risk misuse. Read its dataset card and responsible-use guidance before downloading or using it. ## Selected research from our team - [OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents](https://arxiv.org/abs/2506.14866) — NeurIPS 2025 Spotlight - [AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents](https://arxiv.org/abs/2410.09024) — ICLR 2025 - [JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models](https://arxiv.org/abs/2404.01318) — NeurIPS 2024 - [RobustBench: a Standardized Adversarial Robustness Benchmark](https://arxiv.org/abs/2010.09670) — NeurIPS 2021 ## Work with us We are hiring researchers and engineers. See current roles and learn more about our work at [expsec.ai](https://expsec.ai/#hiring), or contact us at [contact@expsec.ai](mailto:contact@expsec.ai).