# Benchmark: nameplate field extraction One labelled corpus, one metric, so this engine and any other model (Donut, LayoutLMv3, a VLM, a cloud document-AI API) can be compared on the **same** data. ```sh # Run the deterministic engine and write the numbers + the corpus node tests/nameplates-extraction-bench.mjs \ --n=1000 \ --json=ml/nameplates/hf/data-centre-nameplates/eval/extraction-bench.json \ --dump=ml/nameplates/hf/data-centre-nameplates/eval/extraction-corpus.jsonl \ --dumpN=200 ``` ## What it measures Each item is generated from **constructed** field values (so the ground truth is not the parser's own reading), rendered with the wording variety real plates use, and scored field by field: - **clean** — measures the extractor + matcher alone (no OCR loss). - **ocr** — the same plate text pushed through the character errors OCR actually makes on plate lettering (`O/0`, `I/L/1`, `S/5`, `B/8`, `Z/2`, `HZ→Hw`, `KVA→XVA`, dropped/split characters). Reported: per-field precision / recall / F1, whole-plate exact match (no missed and no spurious field), and ms/plate. Code similarity uses the same `O/0`, `I/L/1`, `S/5`, `B/8`, `Z/2` folding the app uses. ## Scoring another model on the same corpus `--dump` writes JSONL, one item per line: ```json {"id":0,"type":"transformer","clean":"","ocr":"","truth":{"manufacturer":"…","model":"…","voltage":[4160,480,277],"kva":1500,…}} ``` Two ways to use it: 1. **Text in / fields out** (LayoutLM, an LLM/VLM over OCR text, a cloud KIE API): feed `ocr`, parse the model's fields, and compare to `truth` with the same normalisers. This isolates extraction. 2. **Image in / fields out** (Donut, TrOCR+KIE, a VLM): render `clean`/`ocr` text to an image with the project's plate drawer (`window.__nameplates.samplePlate` in `nameplates-app.js`, or `ml/nameplates/synth/generate.py`), run the model, compare to `truth`. This folds OCR into the score. Use the project's own comparator so the rules are identical: ```js import { codeSimilarity } from '@constructelligence/nameplate-model'; // or '../nameplate-model.js' const codeOk = (want, got) => codeSimilarity(want, got) >= 0.85; // model / serial const numOk = (want, got) => got != null && Math.abs(+got - +want) < 1e-6; const voltOk = (want, got) => want.values.every(v => got.values.includes(v)); ``` ## Reference points (published, different datasets — not directly comparable) These are reported by their authors on their own benchmarks. They are context for the task, not a head-to-head; the only numbers comparable to each other are ones produced by the corpus above. | Model | Task / dataset | Reported | |---|---|---| | LayoutLMv3 (base) | entity extraction, FUNSD | 92.08 F1 | | LiLT | entity extraction, FUNSD | 88.41 F1 | | Donut | field extraction, CORD | 84.1 F1 | | LayoutLMv3-large | information extraction, FUNSD | ~92 F1 | | CTPN + Transformer | power-equipment nameplate detection / recognition | 88.7% det F1 · 92.3% char acc | | PP-OCRv4 (TL-DREN) | electricity nameplate detection / recognition | 0.524 det F1 · 0.82 rec acc | | RNN nameplate OCR | power-equipment nameplate chars (zh + alnum) | 99.9% zh · 99.3% alnum | Nameplate extraction is a small field with no shared public benchmark — FUNSD/CORD measure invoices and forms, not equipment plates. That is exactly why this corpus is published: to give the task a reproducible, domain-specific test set.