README / README.md
imadreamerboy's picture
Link the ExpSec AI Safety Evaluations collection
fbb7da6 verified
|
Raw
History Blame Contribute Delete
3.2 kB
---
title: README
emoji: 🛡️
colorFrom: indigo
colorTo: red
sdk: static
pinned: false
---
<p align="center">
<img
src="https://huggingface.co/spaces/ExpSec/README/resolve/main/assets/exp-horizontal-cream-rounded-1024.png"
alt="Exponential Security Labs"
width="820"
>
</p>
<p align="center"><strong>Securing the age of agency.</strong></p>
<p align="center">
<a href="https://expsec.ai">Website</a> ·
<a href="https://huggingface.co/ExpSec">Hugging Face</a> ·
<a href="https://huggingface.co/collections/ExpSec/expsec-ai-safety-evaluations-6a5f7308854e355ca281e04c">Collection</a> ·
<a href="https://x.com/expsecai">X</a> ·
<a href="https://www.linkedin.com/company/126363959">LinkedIn</a> ·
<a href="mailto:contact@expsec.ai">Contact</a>
</p>
## About
Exponential Security Labs is an AI safety and security research company. We study AI systems that act through tools and multi-step workflows, and build automated red-teaming and adaptive guardrail methods for them.
Our team brings nearly a decade of work in adversarial robustness and multimodal AI safety. We publish evaluation datasets and develop methods for identifying failures, measuring harmful assistance, and testing safeguards against adaptive attacks.
## What we work on
- **Agent safety evaluation** — benchmarks and evaluation methods for tool-using, multi-step AI systems.
- **Automated red-teaming** — agents and methods that identify vulnerabilities in models, tools, and agent workflows.
- **Adaptive guardrails** — defences designed to respond to changing attacks and deployment contexts.
- **Adversarial robustness** — reliable behaviour under malicious or unexpected inputs across language and vision systems.
## On the Hub
Browse the [ExpSec AI Safety Evaluations collection](https://huggingface.co/collections/ExpSec/expsec-ai-safety-evaluations-6a5f7308854e355ca281e04c) for our public evaluation releases.
### [Sovereign-Jbreak](https://huggingface.co/datasets/ExpSec/Sovereign-Jbreak)
A regional safety evaluation dataset for testing whether a language model gives actionable help towards a harmful objective. The initial release contains expert-authored evaluation items for Europe and India, with goal-specific rubrics and no model responses or attack trajectories.
> **Content warning:** This dataset describes high-risk misuse. Read its dataset card and responsible-use guidance before downloading or using it.
## Selected research from our team
- [OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents](https://arxiv.org/abs/2506.14866) — NeurIPS 2025 Spotlight
- [AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents](https://arxiv.org/abs/2410.09024) — ICLR 2025
- [JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models](https://arxiv.org/abs/2404.01318) — NeurIPS 2024
- [RobustBench: a Standardized Adversarial Robustness Benchmark](https://arxiv.org/abs/2010.09670) — NeurIPS 2021
## Work with us
We are hiring researchers and engineers. See current roles and learn more about our work at [expsec.ai](https://expsec.ai/#hiring), or contact us at [contact@expsec.ai](mailto:contact@expsec.ai).