Instructions to use HarshilDaGoat/gem-certificate-extractor with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use HarshilDaGoat/gem-certificate-extractor with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="HarshilDaGoat/gem-certificate-extractor")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("HarshilDaGoat/gem-certificate-extractor") model = AutoModelForTokenClassification.from_pretrained("HarshilDaGoat/gem-certificate-extractor", device_map="auto") - Notebooks
- Google Colab
- Kaggle
GeM certificate field extractor
Token-classification model that reads the OCR text of Indian business certificates submitted with Government e-Marketplace (GeM) bids and extracts the fields a compliance officer cross-checks: GSTIN, PAN, Udyam number, CIN, EPFO establishment code, DPIIT number, legal and trade name, activity, enterprise category, certified local-content % and validity-until date.
Part of an automated bid-compliance platform; it replaces a regex parser in the document stage, and its output is cross-verified against government portal data by an auditable rule engine.
Intended use
Input: plain text from a certificate (PDF text layer or OCR). Output: one value per field, with a confidence. Covered document types: GST registration certificate (REG-06), Udyam certificate, Make in India local-content self-certificate, PAN card, EPFO compliance certificate, DPIIT startup recognition, OEM authorization letter, CA turnover certificate, NSIC registration.
It is an extraction aid, not a verifier: a correctly extracted GSTIN says nothing about whether that GSTIN is active โ that comes from the GST portal.
Training data
Generated by ml/extraction/generate.py: 54,000 train, 3,000 validation, 3,000 test examples.
The generator deliberately includes look-alikes a regex gets wrong: the OEM's own GSTIN beside the bidder's in authorization letters, a "minimum local content of N%" threshold restated before the certified figure, validity start dates, PANs embedded in GSTINs, past-year Udyam classifications, and UDIN/FRN/ESIC numbers โ plus OCR-style character confusions (O/0, S/5, I/1) and merged lines.
All training data is synthetic. Identifiers follow real formats (GSTINs carry a valid check character) but correspond to no real entity. Accuracy on real, scanned certificates with layouts unlike the templates will be lower than the figures below; evaluate on your own documents first.
Evaluation (held-out synthetic test set)
Entity-level, exact span match:
| Field | Precision | Recall | F1 | Support |
|---|---|---|---|---|
| ACTIVITY | 100.0% | 100.0% | 100.0% | 902 |
| CATEGORY | 100.0% | 100.0% | 100.0% | 902 |
| CIN | 100.0% | 100.0% | 100.0% | 124 |
| DPIIT | 100.0% | 100.0% | 100.0% | 312 |
| EPFO | 100.0% | 100.0% | 100.0% | 311 |
| GSTIN | 100.0% | 100.0% | 100.0% | 1174 |
| LEGAL_NAME | 100.0% | 100.0% | 100.0% | 3000 |
| LOCAL_CONTENT | 100.0% | 99.7% | 99.8% | 296 |
| PAN | 100.0% | 100.0% | 100.0% | 593 |
| TRADE_NAME | 100.0% | 100.0% | 100.0% | 296 |
| UDYAM | 100.0% | 100.0% | 100.0% | 902 |
| VALID_UNTIL | 100.0% | 100.0% | 100.0% | 789 |
| All (micro) | 100.0% | 100.0% | 100.0% | 9601 |
Usage
from transformers import pipeline
extract = pipeline("token-classification", model="HarshilDaGoat/gem-certificate-extractor", aggregation_strategy="simple")
extract("Registration Number: 27AABCN1234A1Z5\nLegal Name of Business: Nova Electro Systems Private Limited")
For field values normalised exactly as the platform uses them (dates as ISO, local content as a
number, one value per field), use ml/extraction/infer.py from the project repository.
- Downloads last month
- 86
Model tree for HarshilDaGoat/gem-certificate-extractor
Base model
distilbert/distilbert-base-cased