Spaces:
Running
Running
| title: README | |
| emoji: 🛡️ | |
| colorFrom: indigo | |
| colorTo: red | |
| sdk: static | |
| pinned: false | |
| <p align="center"> | |
| <img | |
| src="https://huggingface.co/spaces/ExpSec/README/resolve/main/assets/exp-horizontal-cream-rounded-1024.png" | |
| alt="Exponential Security Labs" | |
| width="820" | |
| > | |
| </p> | |
| <p align="center"><strong>Securing the age of agency.</strong></p> | |
| <p align="center"> | |
| <a href="https://expsec.ai">Website</a> · | |
| <a href="https://huggingface.co/ExpSec">Hugging Face</a> · | |
| <a href="https://huggingface.co/collections/ExpSec/expsec-ai-safety-evaluations-6a5f7308854e355ca281e04c">Collection</a> · | |
| <a href="https://x.com/expsecai">X</a> · | |
| <a href="https://www.linkedin.com/company/126363959">LinkedIn</a> · | |
| <a href="mailto:contact@expsec.ai">Contact</a> | |
| </p> | |
| ## About | |
| Exponential Security Labs is an AI safety and security research company. We study AI systems that act through tools and multi-step workflows, and build automated red-teaming and adaptive guardrail methods for them. | |
| Our team brings nearly a decade of work in adversarial robustness and multimodal AI safety. We publish evaluation datasets and develop methods for identifying failures, measuring harmful assistance, and testing safeguards against adaptive attacks. | |
| ## What we work on | |
| - **Agent safety evaluation** — benchmarks and evaluation methods for tool-using, multi-step AI systems. | |
| - **Automated red-teaming** — agents and methods that identify vulnerabilities in models, tools, and agent workflows. | |
| - **Adaptive guardrails** — defences designed to respond to changing attacks and deployment contexts. | |
| - **Adversarial robustness** — reliable behaviour under malicious or unexpected inputs across language and vision systems. | |
| ## On the Hub | |
| Browse the [ExpSec AI Safety Evaluations collection](https://huggingface.co/collections/ExpSec/expsec-ai-safety-evaluations-6a5f7308854e355ca281e04c) for our public evaluation releases. | |
| ### [Sovereign-Jbreak](https://huggingface.co/datasets/ExpSec/Sovereign-Jbreak) | |
| A regional safety evaluation dataset for testing whether a language model gives actionable help towards a harmful objective. The initial release contains expert-authored evaluation items for Europe and India, with goal-specific rubrics and no model responses or attack trajectories. | |
| > **Content warning:** This dataset describes high-risk misuse. Read its dataset card and responsible-use guidance before downloading or using it. | |
| ## Selected research from our team | |
| - [OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents](https://arxiv.org/abs/2506.14866) — NeurIPS 2025 Spotlight | |
| - [AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents](https://arxiv.org/abs/2410.09024) — ICLR 2025 | |
| - [JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models](https://arxiv.org/abs/2404.01318) — NeurIPS 2024 | |
| - [RobustBench: a Standardized Adversarial Robustness Benchmark](https://arxiv.org/abs/2010.09670) — NeurIPS 2021 | |
| ## Work with us | |
| We are hiring researchers and engineers. See current roles and learn more about our work at [expsec.ai](https://expsec.ai/#hiring), or contact us at [contact@expsec.ai](mailto:contact@expsec.ai). | |