GeM certificate field extractor

Token-classification model that reads the OCR text of Indian business certificates submitted with Government e-Marketplace (GeM) bids and extracts the fields a compliance officer cross-checks: GSTIN, PAN, Udyam number, CIN, EPFO establishment code, DPIIT number, legal and trade name, activity, enterprise category, certified local-content % and validity-until date.

Part of an automated bid-compliance platform; it replaces a regex parser in the document stage, and its output is cross-verified against government portal data by an auditable rule engine.

Intended use

Input: plain text from a certificate (PDF text layer or OCR). Output: one value per field, with a confidence. Covered document types: GST registration certificate (REG-06), Udyam certificate, Make in India local-content self-certificate, PAN card, EPFO compliance certificate, DPIIT startup recognition, OEM authorization letter, CA turnover certificate, NSIC registration.

It is an extraction aid, not a verifier: a correctly extracted GSTIN says nothing about whether that GSTIN is active โ€” that comes from the GST portal.

Training data

Generated by ml/extraction/generate.py: 54,000 train, 3,000 validation, 3,000 test examples.

The generator deliberately includes look-alikes a regex gets wrong: the OEM's own GSTIN beside the bidder's in authorization letters, a "minimum local content of N%" threshold restated before the certified figure, validity start dates, PANs embedded in GSTINs, past-year Udyam classifications, and UDIN/FRN/ESIC numbers โ€” plus OCR-style character confusions (O/0, S/5, I/1) and merged lines.

All training data is synthetic. Identifiers follow real formats (GSTINs carry a valid check character) but correspond to no real entity. Accuracy on real, scanned certificates with layouts unlike the templates will be lower than the figures below; evaluate on your own documents first.

Evaluation (held-out synthetic test set)

Entity-level, exact span match:

Field Precision Recall F1 Support
ACTIVITY 100.0% 100.0% 100.0% 902
CATEGORY 100.0% 100.0% 100.0% 902
CIN 100.0% 100.0% 100.0% 124
DPIIT 100.0% 100.0% 100.0% 312
EPFO 100.0% 100.0% 100.0% 311
GSTIN 100.0% 100.0% 100.0% 1174
LEGAL_NAME 100.0% 100.0% 100.0% 3000
LOCAL_CONTENT 100.0% 99.7% 99.8% 296
PAN 100.0% 100.0% 100.0% 593
TRADE_NAME 100.0% 100.0% 100.0% 296
UDYAM 100.0% 100.0% 100.0% 902
VALID_UNTIL 100.0% 100.0% 100.0% 789
All (micro) 100.0% 100.0% 100.0% 9601

Usage

from transformers import pipeline
extract = pipeline("token-classification", model="HarshilDaGoat/gem-certificate-extractor", aggregation_strategy="simple")
extract("Registration Number: 27AABCN1234A1Z5\nLegal Name of Business: Nova Electro Systems Private Limited")

For field values normalised exactly as the platform uses them (dates as ISO, local content as a number, one value per field), use ml/extraction/infer.py from the project repository.

Downloads last month
86
Safetensors
Model size
65.2M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for HarshilDaGoat/gem-certificate-extractor

Finetuned
(339)
this model