AI & ML interests
Evaluation benchmarks and training datasets
Recent Activity
An AI capability and efficiency lab for teams building and improving frontier systems.
We produce expert-generated data and research-grade evaluation for frontier model
development — and we publish our benchmarks openly rather than keeping them as sales material.
What we do
Frontier model training
Expert-generated data pipelines across pre-training, fine-tuning and reinforcement learning, built to each model's specification.
Agent evaluation & optimisation
Structured evaluation and human-in-the-loop refinement, taking autonomous agents from prototype to production-grade reliability.
Red teaming and safety
Adversarial testing and safety evaluation by trained specialists, finding vulnerabilities before users do.
Vertical and STEM expertise
Domain-expert data and evaluation for the fields where generic training data falls short.
Open benchmarks
PDL-Bench
98 specialist tasks across six mini-benchmarks, written by our experts to beat frontier models. Every task ships with what its grading needs — a final answer, a rubric, a worked solution, or an executable test suite — so results reproduce locally.
| Mini-benchmark | Focus | Tasks |
|---|---|---|
| Pure Mathematics | Proof, construction, exact quantitative reasoning | 16 |
| Advanced Science | Research-level scientific interpretation and mechanism | 39 |
| Business Operations | Operational analysis and professional deliverables | 9 |
| Life Administration | Everyday planning, research and personal logistics | 4 |
| Software Delivery | Repository-scale engineering in executable environments | 20 |
| Cyber Operations | Security-critical protocol and trust-system implementation | 10 |
PDL-SWE-Bench
18 repository-scale software engineering tasks in SWE-bench schema, with hidden test suites and gold patches.
Both are CC BY 4.0. Where a suite shows a model is weak, that is exactly the training data we build.
The people who do the work
Poindexters are STEM domain experts, selected through subject-knowledge interviews for the depth and judgment that frontier AI work demands.