--- license: apache-2.0 language: - en library_name: custom pipeline_tag: image-to-text tags: - ocr - document-ai - information-extraction - key-information-extraction - nameplate - nameplate-ocr - equipment-nameplate - data-center - data-centre - datacenter - mission-critical - mep - electrical - mechanical - equipment-schedule - schedule-verification - asset-register - commissioning - quality-assurance - ups - pdu - switchgear - generator - ats - transformer - busway - battery - crah - crac - chiller - cooling - server-rack - construction - aec - edge-ai - on-device - privacy-preserving - tesseract - browser metrics: - accuracy - f1 model-index: - name: data-centre-nameplates results: - task: type: image-to-text name: Equipment-nameplate field extraction & schedule verification dataset: name: Data Centre Construction labelled plate corpus (n = 1000) type: constructelligence/nameplate-corpus split: clean metrics: - type: accuracy value: 0.941 name: Exact-plate accuracy (clean) - type: f1 value: 0.997 name: Micro-F1 (clean) - type: f1 value: 0.799 name: Micro-F1 (character-error stress) --- # Data Centre Construction — equipment-nameplate OCR checked against the equipment schedule > **Constructelligence is developing frontier construction-AI models.** This kit is the *lite*, open tier — > released for research and evaluation; the company's flagship production models are a separate, larger > tier (see [constructelligence.co](https://constructelligence.co)). **Read an equipment nameplate and know if it is the unit the schedule asked for.** Point a phone at the nameplate of a UPS, PDU, switchgear, generator, transformer, busway, battery, CRAH, chiller — or a **server rack** — and the engine returns the plate's fields (35 of them, grouped Identification / Electrical / Cooling / Server racks / Mechanical) and **checks each one against the equipment schedule**, flagging a mismatch instead of a silent pass. It builds an asset register and exports it to CSV. Built for the part of a data-centre job where the gear *is* the job. The expensive mistakes are a unit delivered at the wrong voltage or rating, or one that quietly never makes it onto the asset register. > **Runs on the device.** OCR, extraction and matching run in the browser (or Node) with no server and no > upload. A nameplate photo of live infrastructure never leaves the phone that took it — no GPU, no key, no > data egress. **▶ Try it live:** [huggingface.co/spaces/constructelligence/data-centre-nameplates](https://huggingface.co/spaces/constructelligence/data-centre-nameplates) — a browser demo with a sample schedule and sample plates. | | | |---|---| | **Task** | Image → structured nameplate fields → schedule verification | | **Fields** | **35**, in five groups (Identification, Electrical, Cooling, Server racks/IT, Mechanical) | | **Equipment** | **55 types**, device wording ranked above family (see below) | | **OCR** | Tesseract.js 5 (English LSTM), pretrained — no fine-tuning | | **Extraction** | Deterministic labelled-field parser + OCR error model + numeric range guards | | **Matching** | Code folding (`O/0`, `I/L/1`, `S/5`, `B/8`, `Z/2`) + voltage-compatibility signal | | **Quality gate** | Advisory photo check (blur, glare, darkness, resolution) before OCR | | **Runtime** | Browser · Node · edge — no GPU, no server, no data egress | | **Field accuracy** | **99.8%** clean / **80.0%** under character-error stress (extraction benchmark) | | **Licence** | Apache-2.0 | --- ## Why this exists General OCR reads the text on a nameplate. A commissioning or QA team needs the text to answer a narrower question: **is this the right unit, and does it match what was specified?** That means three things a plain OCR call does not do: 1. **Turn plate text into fields, not paragraphs.** "INPUT: 480Y/277 VAC 3 PH 60 HZ" becomes `voltage=480Y/277 V`, `phase=3`, `hz=60`; "RATING: 750 KVA / 675 KW" becomes `kva=750`, `kw=675`. 2. **Survive the way OCR misreads plate lettering.** Serifed `V`→`Vv`, `K`→`X` (`XVA`), `HZ`→`Hw`, codes split after their punctuation (`NPX- 750- 480`). The parser repairs these before matching, and rejects out-of-range noise (`0 KVA`, a stray `0` amp) rather than recording it. 3. **Compare to the schedule the way codes actually differ.** Model and serial codes are compared with `O/0`, `I/L/1`, `S/5`, `B/8`, `Z/2` folded and punctuation ignored, so a misread is not a false mismatch, while a genuine voltage variant (`NPX-750-415` vs `NPX-750-480`) still fails. A voltage a plate and a schedule row share is a positive signal; the equipment **type** still wins ties, so a 415 V UPS plate is not matched to a same-voltage PDU. The result is a pass/mismatch verdict per plate, with the fields that disagree named. ## How it works ``` photo ─► photo check ─► OCR (Tesseract.js) ─► field parser ─► equipment typing ─► schedule matcher ─► verdict + CSV blur/glare lines + confidence labelled rules maker line excluded code folding, voltage signal ``` 1. **Photo check** — an advisory blur/glare/darkness/resolution read before OCR; the shot can be retaken before a bad read. 2. **OCR** (`tesseract.js@5`, English) returns lines with a per-line confidence; two page-segmentation passes are merged field by field when the first is weak. 3. **Field parser** scans lines for labelled values (`MODEL`, `S/N`, `MVA`, `MCA`, `MOCP`, `FLA`, `VOLTS`, `REFRIGERANT`, `MFG DATE`, …), reading voltages as `480Y/277`, `13.8 kV`, `208/120`, and rejecting look-alikes. Numeric fields are range-guarded. Lines OCR scored near zero are dropped first. 4. **Equipment typing** recognises **55 families** and ranks the specific device above its family (flywheel before UPS; paralleling gear before generator; dry cooler before chiller; rack before server). The **maker line is excluded** from typing, so a vendor called “…Transformer Corp” cannot type a chiller plate. 5. **Schedule matcher** ranks candidate tags — a serial match wins outright; otherwise model similarity dominates, with manufacturer, equipment type and a compatible voltage as signals. Among equal matches the first **unscanned** tag is offered. 6. **Field checks** compare the reading to the chosen row (`ok` / `mismatch` / `unread` per field) and return an overall status: `verified`, `partial`, `mismatch`, `unmatched` or `missing` (scheduled but not scanned). ## Fields extracted (35) **Identification** — `manufacturer`, `model`, `serial`, `mfgDate` **Electrical** | Field | Example | Notes | |---|---|---| | `voltage` | 480Y/277 V | every voltage on the plate; `13.8 kV`, `208/120`, `480` | | `phase` / `hz` | 3 / 60 | `PH`/`PHASE`/`Ø`; 50 or 60 Hz | | `kva` / `kw` | 750 / 675 | `kVA`, `KW`, `EKW`; thousands separators | | `amps` | 1002 | labelled `FLA`/`RLA`/`CURRENT`/`AMPS` first | | `pf` | 0.8 | power factor, 0–1 | | `mca` / `mocp` | 612 / 800 | min circuit ampacity; max overcurrent protection | | `sccr` | 65 | short-circuit / interrupting rating (kAIC) | | `enclosure` | NEMA 3R | NEMA type or IEC IP code | | `standards` | UL 1989 | UL / IEC / ANSI / IEEE / NFPA / CSA / EN listings | | `batteryType` | VRLA | VRLA, AGM, Gel, Flooded, Li-ion, LiFePO4, NiCd, Lead-acid | | `batteryAh` / `cells` | 500 / 48 | amp-hours; cell count | **Cooling** — `refrigerant`, `charge`, `tons`, `btu`, `cfm`, `gpm` **Server racks / IT** — `uHeight` (rack U, incl. "42RU"), `formFactor` (blade / tower / a labelled "2U"), `powerW`, `rackLoad` (static/dynamic load, kg), `outletTypes` (canonical outlet mix, e.g. `C13x12 C19x4`), `outlets`, `portSpeed` (fastest link, GbE), `ports` **Mechanical** — `rpm`, `weight` (normalised to kg) ## Equipment types recognised (55) Power: flywheel · UPS · STS · ATS · rectifier/DC · inverter · BESS · nitrogen generator · paralleling gear · generator · fuel system · switchgear/switchboard · MCC · metering/CT · protective relay · transformer/DGA monitor · SPD · capacitor/PFC · neutral grounding resistor · VFD · busway · panelboard · PDU/RPP · battery · load bank Cooling & mechanical: heat pump · pump · cooling tower · thermal storage · chiller · dry cooler/condenser · CDU · CRAH/CRAC · AHU · heat exchanger · expansion tank · air separator · boiler · humidifier · water treatment · air dryer · compressor · fan coil · rooftop unit · makeup air · economizer Fire, safety & controls: fire suppression · fire alarm · leak/gas detection · BMS/EPMS/DCIM IT & white space: **rack** · **server/storage** · network gear · KVM/console ## Server racks — a first-class subsection Racks, rack PDUs, switches and in-row cooling are what the electrical and cooling design actually serves, so they are treated as data-centre equipment, not an afterthought: `Rack U`, `Form factor`, `Power (W)`, `Rack load (kg)`, `Outlet mix`, `PDU outlets`, `Port speed (GbE)` and `Ports` are extracted and checked like any other field, and the equipment set covers rack/enclosure, server, network, KVM and rack PDU. Sample plates include a PowerEdge-class server, a rack PDU with a C13/C19 outlet mix, a 25G switch and a blade chassis with a static floor load. ## Results ### Extraction benchmark — labelled plates, field by field A generated corpus of plates with **by-construction** ground truth (not the parser's own reading), rendered with wording variety. `clean` isolates the extractor + matcher; `ocr` adds the character errors OCR makes on plate lettering. n = 1000. Reproduce with `node tests/nameplates-extraction-bench.mjs --n=1000`. | Condition | Exact plate | Micro-F1 | ms/plate | |---|:---:|:---:|---:| | clean | **94.1%** | **99.7%** | 0.49 | | ocr (character-error stress) | 1.1% | **79.9%** | 0.49 | Per-field F1 is 100% for manufacturer, type, model, serial, kVA, phase, Hz, MCA, MOCP, refrigerant, pf, cfm, uHeight, sccr and enclosure in the clean condition; the residual clean misses are a small number of voltage and `mfgDate` cases. The `ocr` condition is deliberately harsh (per-character flips); low-confidence fields are surfaced for a human rather than trusted. Per-field tables: `eval/extraction-bench.json`. ### OCR benchmark — drawn plates through phone-photo degradations Plates drawn for several equipment types, each pushed through real-photo degradations (small, blurred, noisy, low-contrast, glare, tilted, keystoned, dark anodised, JPEG, plate-small-in-scene, sideways, and a "phone" composite), then run through the app's own OCR + parser and scored field by field. An earlier run over six plate types and 13 conditions scored **99.9% overall (727/728)**; the harness now draws more plate types and is re-run on release. Run it with `node tests/nameplates-bench.mjs` (needs a browser for Tesseract.js). **Read this honestly.** These are machine-drawn plates. A real plate on a live unit is harder: cast shadows, embossed lettering, print over brushed metal, a plate half out of frame. Expect the degradation suite, not the clean row, and treat every low-confidence field as one a person must confirm. ## Comparison with related models Nameplate extraction has no shared public benchmark — FUNSD/CORD measure invoices and forms, not equipment plates. These are the closest published reference points, on **their own** datasets; they are context, not a head-to-head. The only directly comparable scores are ones produced by the corpus in `ml/nameplates/bench/`. | Model / method | Task & dataset | Reported | |---|---|---| | **This engine** | **named-field extraction, own labelled plate corpus** | **99.7% micro-F1 (clean)** | | LayoutLMv3 (base) | entity extraction, FUNSD | 92.08 F1 | | LiLT | entity extraction, FUNSD | 88.41 F1 | | Donut | field extraction, CORD | 84.1 F1 | | CTPN + Transformer | power-equipment nameplate detection / recognition | 88.7% det F1 · 92.3% char acc | | PP-OCRv4 (TL-DREN) | electricity nameplate detection / recognition | 0.524 det F1 · 0.82 rec acc | | RNN nameplate OCR | power-equipment nameplate chars (zh + alnum) | 99.9% zh · 99.3% alnum | | OCR-free VLM (e.g. Qwen2.5-VL-7B) | document KIE, zero-shot | varies; strong but large | What is different here: the engine is **small, deterministic, on-device, and schedule-aware**. It does not just transcribe a plate — it decides `verified` / `mismatch` against a schedule and names the fields that disagree, which no general document model does out of the box. ## Fine-tuning — the higher-accuracy server model Alongside the on-device path, a **Donut** model ([`naver-clova-ix/donut-base`](https://huggingface.co/naver-clova-ix/donut-base)) is fine-tuned to read a plate **straight into fields** (`ups…`) for a server / endpoint where the browser engine falls short. Training data is synthetic plates with real-photo degradations plus corrected real plates; best checkpoints are selected by **field accuracy**, using the same OCR-aware comparison the app uses. The corpus in `ml/nameplates/bench/` scores it on the same test set as the on-device engine. Nothing on this page claims the accuracy of that model yet — it is the roadmap, not a result. ## Intended use - **Commissioning and QA walks:** photograph each unit as it is installed; confirm the delivered unit matches the equipment schedule; capture an asset register as you go. - **Receiving and delivery checks:** flag the unit that arrived at the wrong voltage or rating before it is set. - **Register completion:** track scheduled tags that have no plate scanned against them yet. - **Takeoff and submittal review:** pull model/rating fields off a plate photo into a spreadsheet. ### Out of scope — do not use it for - **Safety, code compliance or energisation decisions.** It reads a plate; it does not verify that a unit is safe to energise, correctly protected, or code-compliant. - **An as-built or a legal record on its own.** OCR misreads happen; low-confidence fields are highlighted and any field that disagrees with the schedule must be confirmed by eye. - **Serial-number evidence in a dispute** without human verification; the register flags duplicate serials but cannot tell a repeated plate from a copied number. ## Limitations - **English nameplates**, Latin script. Non-Latin or handwritten plates are out of scope. - **Glare, blur and angle degrade OCR.** The parser repairs common misreads, but a badly glared or very low-contrast plate will lose fields (those come back `unread`, not invented). - **Deterministic parser, not a language model.** It reports only what it can match to a rule; it will not infer a missing digit. - **The schedule must be structured** (CSV with a Tag or Model column). Free-form schedule PDFs are not parsed here. ## Quickstart The engine is plain ES modules with no build step. `nameplate-model.js` is the extractor and verifier; `schedule-import.js` is the CSV reader it depends on. ```js import { parseNameplate, parseEquipmentSchedule, candidatesFor, checkAgainst, statusOf } from './nameplate-model.js'; const reading = parseNameplate(`NORTHLINE POWER SYSTEMS UNINTERRUPTIBLE POWER SUPPLY MODEL NO: NPX-750-480 SERIAL NO: NL26A01937 INPUT: 480Y/277 VAC 3 PH 60 HZ RATING: 750 KVA / 675 KW INPUT CURRENT: 1002 A MFG DATE: 03/2026`); console.log(reading.type); // 'ups' console.log(reading.fields.voltage); // { value: '480Y/277 V', values: [480, 277], conf, line } const { items } = parseEquipmentSchedule(scheduleCsv); const [best] = candidatesFor(reading, items); console.log(statusOf(checkAgainst(reading, best.item), best.item.tag)); // 'verified' | 'partial' | 'mismatch' | … ``` In a browser, feed Tesseract's `data.lines` (`[{ text, confidence, bbox }]`) straight into `parseNameplate` to keep OCR confidence and the plate line each value came from. ## Evaluation & reproducibility - **Extraction:** `node tests/nameplates-extraction-bench.mjs --n=1000 --json=…` — labelled corpus, per-field precision/recall/F1, whole-plate exact match, ms/plate. Corpus dump for other models: `… --dump=…/extraction-corpus.jsonl`. Method and comparator: `ml/nameplates/bench/README.md`. - **Unit tests:** the parser, voltage reader, code folding, schedule parsing, candidate ranking, mismatch detection, range guards and the matcher — `node --test tests/nameplate-model.test.mjs` (all passing). - **OCR end-to-end:** `node tests/nameplates-bench.mjs` — drawn plates through phone-photo degradations. - **Fine-tuned reader:** `ml/nameplates/cloud/eval_donut.py` scores the Donut model field by field on the held-out split. ## Repository contents | File | What it is | |---|---| | `nameplate-model.js` | the extractor, matcher and register builder (ES module) | | `schedule-import.js` | the CSV schedule reader it depends on | | `progress-project.js` | a helper the schedule reader imports | | `sample_schedule.csv` | a 10-row data-centre equipment schedule | | `sample_plates.json` | sample plates: a match, a flagged mismatch, a switchboard, a server | | `eval/` | OCR benchmark output, extraction benchmark output and corpus | | `config.json` | machine-readable fields, groups, types and statuses | ## Provenance An open extraction and verification engine from Constructelligence, the same code path that powers the **Data Centre Construction** mode of BuildVision. No customer data is used; sample plates use fictional makers and models. It is the document-AI sibling of the vision models [`construction-site-safety-hazards`](https://huggingface.co/constructelligence/construction-site-safety-hazards) and [`electrical-circuit-connectivity`](https://huggingface.co/constructelligence/electrical-circuit-connectivity). ## Safety This is a **prompt to look**, not a finding. It is not safety-rated, not a compliance decision, and not a substitute for inspection by a competent person. Verify every flagged field against the physical plate before acting on it. ## Licence and trademarks Code released under **Apache-2.0**. Brand and product names (Schneider Electric, Eaton, Vertiv, Caterpillar, Trane, Dell, Cisco, …) are trademarks of their owners and are matched only to identify user-supplied plates; this project is independent and not affiliated with or endorsed by them. ## Citation ```bibtex @misc{constructelligence_nameplates, title = {Data Centre Construction: equipment-nameplate OCR checked against the equipment schedule}, author = {Constructelligence}, year = {2026}, howpublished = {\url{https://huggingface.co/constructelligence/data-centre-nameplates}} } ```