AI & ML interests

Evaluation benchmarks and training datasets

Recent Activity

Organization Card

An AI capability and efficiency lab for teams building and improving frontier systems.
We produce expert-generated data and research-grade evaluation for frontier model development — and we publish our benchmarks openly rather than keeping them as sales material.

What we do

Frontier model training

Expert-generated data pipelines across pre-training, fine-tuning and reinforcement learning, built to each model's specification.

Agent evaluation & optimisation

Structured evaluation and human-in-the-loop refinement, taking autonomous agents from prototype to production-grade reliability.

Red teaming and safety

Adversarial testing and safety evaluation by trained specialists, finding vulnerabilities before users do.

Vertical and STEM expertise

Domain-expert data and evaluation for the fields where generic training data falls short.

Open benchmarks

PDL-Bench

98 specialist tasks across six mini-benchmarks, written by our experts to beat frontier models. Every task ships with what its grading needs — a final answer, a rubric, a worked solution, or an executable test suite — so results reproduce locally.

Mini-benchmark Focus Tasks
Pure Mathematics Proof, construction, exact quantitative reasoning 16
Advanced Science Research-level scientific interpretation and mechanism 39
Business Operations Operational analysis and professional deliverables 9
Life Administration Everyday planning, research and personal logistics 4
Software Delivery Repository-scale engineering in executable environments 20
Cyber Operations Security-critical protocol and trust-system implementation 10

PDL-SWE-Bench

18 repository-scale software engineering tasks in SWE-bench schema, with hidden test suites and gold patches.

Both are CC BY 4.0. Where a suite shows a model is weak, that is exactly the training data we build.

The people who do the work

Poindexters are STEM domain experts, selected through subject-knowledge interviews for the depth and judgment that frontier AI work demands.

46%Olympiad medallists
54%PhD level
34%Oxbridge, Ivy League or MIT
18%Professor level

models 0

None public yet