Pivot / docs /PERFORMANCE.md
Q1z's picture
Expand Pivot model card, benchmarks, CPU tools and charts
7c85c7e verified
|
Raw
History Blame Contribute Delete
1.87 kB
# Pivot performance
All figures here refer to checkpoint `14bf8c26bf344ebdf88e22a4b6152dc5f75f3578`, public JevBench v1.4.1 commit `24b9b5c1609a7a9e8fa14f49e5985a836c9dc842`, FP32 and the same frozen 512/128-token input contract.
## Accuracy on public tasks
| Tier | Correct | Tasks | Accuracy | Top-label ECE, 10 bins |
|---|---:|---:|---:|---:|
| Original | 27 | 72 | 37.50% | 0.5247 |
| Easy | 39 | 48 | 81.25% | 0.1139 |
| Hard | 41 | 111 | 36.94% | 0.3792 |
| **Total** | **107** | **231** | **46.32%** | — |
![Public-tier accuracy](../evaluation/2026-09-24/accuracy.png)
The official JevBench composite score is **unavailable** because the sealed and judge tasks and official cost input were not measured.
## Local speed
| Warm local FP32 measure | H200 GPU | Xeon CPU, 4 threads |
|---|---:|---:|
| Single decision p50, 32 measured | 15.7668 ms | 797.5558 ms |
| Single decision p95, 32 measured | 19.8708 ms | 1089.0266 ms |
| Single decision mean | 16.1614 ms | 770.7159 ms |
| Batch size for throughput | 32 | 4 |
| Throughput, median of 3 × 64 decisions | 545.2833 decisions/s | 3.7708 decisions/s |
![Warm local latency and throughput](../evaluation/2026-09-24/latency_throughput.png)
The CPU p50 single-decision time is **50.6×** the H200 p50 for these two machines. CPU and GPU batch throughput used different batch sizes and should not be read as a same-batch comparison. The timings cover tokenizer + inference + scoring with 5 warmup singles and 2 warmup bulk passes. Reproduce them on your own hardware using the [CPU script](../cpu-speed/README.md) or [full public runner](../benchmarks/README.md).
Source: [full structured summary](../evaluation/2026-09-24/performance.json), [original public benchmark result](../evaluation/2026-09-24/jevbench_public.json), and [original CPU measurement](../evaluation/2026-09-24/cpu_speed.json).