constructelligence's picture
Update data-centre nameplates to v2: 31 fields, 55 equipment types, extraction benchmark
2843cff verified
|
Raw History Blame Contribute Delete
3.48 kB

Benchmark: nameplate field extraction

One labelled corpus, one metric, so this engine and any other model (Donut, LayoutLMv3, a VLM, a cloud document-AI API) can be compared on the same data.

# Run the deterministic engine and write the numbers + the corpus
node tests/nameplates-extraction-bench.mjs \
  --n=1000 \
  --json=ml/nameplates/hf/data-centre-nameplates/eval/extraction-bench.json \
  --dump=ml/nameplates/hf/data-centre-nameplates/eval/extraction-corpus.jsonl \
  --dumpN=200

What it measures

Each item is generated from constructed field values (so the ground truth is not the parser's own reading), rendered with the wording variety real plates use, and scored field by field:

  • clean — measures the extractor + matcher alone (no OCR loss).
  • ocr — the same plate text pushed through the character errors OCR actually makes on plate lettering (O/0, I/L/1, S/5, B/8, Z/2, HZ→Hw, KVA→XVA, dropped/split characters).

Reported: per-field precision / recall / F1, whole-plate exact match (no missed and no spurious field), and ms/plate. Code similarity uses the same O/0, I/L/1, S/5, B/8, Z/2 folding the app uses.

Scoring another model on the same corpus

--dump writes JSONL, one item per line:

{"id":0,"type":"transformer","clean":"<plate text>","ocr":"<corrupted text>","truth":{"manufacturer":"…","model":"…","voltage":[4160,480,277],"kva":1500,…}}

Two ways to use it:

  1. Text in / fields out (LayoutLM, an LLM/VLM over OCR text, a cloud KIE API): feed ocr, parse the model's fields, and compare to truth with the same normalisers. This isolates extraction.
  2. Image in / fields out (Donut, TrOCR+KIE, a VLM): render clean/ocr text to an image with the project's plate drawer (window.__nameplates.samplePlate in nameplates-app.js, or ml/nameplates/synth/generate.py), run the model, compare to truth. This folds OCR into the score.

Use the project's own comparator so the rules are identical:

import { codeSimilarity } from '@constructelligence/nameplate-model'; // or '../nameplate-model.js'
const codeOk = (want, got) => codeSimilarity(want, got) >= 0.85;        // model / serial
const numOk  = (want, got) => got != null && Math.abs(+got - +want) < 1e-6;
const voltOk = (want, got) => want.values.every(v => got.values.includes(v));

Reference points (published, different datasets — not directly comparable)

These are reported by their authors on their own benchmarks. They are context for the task, not a head-to-head; the only numbers comparable to each other are ones produced by the corpus above.

Model Task / dataset Reported
LayoutLMv3 (base) entity extraction, FUNSD 92.08 F1
LiLT entity extraction, FUNSD 88.41 F1
Donut field extraction, CORD 84.1 F1
LayoutLMv3-large information extraction, FUNSD ~92 F1
CTPN + Transformer power-equipment nameplate detection / recognition 88.7% det F1 · 92.3% char acc
PP-OCRv4 (TL-DREN) electricity nameplate detection / recognition 0.524 det F1 · 0.82 rec acc
RNN nameplate OCR power-equipment nameplate chars (zh + alnum) 99.9% zh · 99.3% alnum

Nameplate extraction is a small field with no shared public benchmark — FUNSD/CORD measure invoices and forms, not equipment plates. That is exactly why this corpus is published: to give the task a reproducible, domain-specific test set.