Image-Text-to-Text
Safetensors
MLX
mlx-vlm
indic_ocr
ocr
document-parsing
layout-analysis
reading-order
indic
Instructions to use HashNuke/indic-ocr-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use HashNuke/indic-ocr-mlx with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("HashNuke/indic-ocr-mlx") config = load_config("HashNuke/indic-ocr-mlx") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Clarify IndicOCR usage and stage loading
Browse filesDocument page parsing and standalone stage inference. Correct model metadata, license links, and validation scope.
- README.md +87 -15
- weights/layout/README.md +32 -10
- weights/ocr/README.md +42 -24
README.md
CHANGED
|
@@ -3,36 +3,69 @@ language:
|
|
| 3 |
- en
|
| 4 |
- as
|
| 5 |
- bn
|
|
|
|
|
|
|
|
|
|
| 6 |
- hi
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 7 |
- mr
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
- ta
|
| 9 |
- te
|
|
|
|
| 10 |
tags:
|
| 11 |
- mlx
|
| 12 |
- ocr
|
| 13 |
- document-parsing
|
|
|
|
|
|
|
| 14 |
- indic
|
| 15 |
library_name: mlx-vlm
|
| 16 |
pipeline_tag: image-text-to-text
|
| 17 |
license: other
|
| 18 |
license_name: indic-open-model-license-1.0
|
| 19 |
-
license_link: LICENSE
|
| 20 |
base_model: bodhan-ai/indic-ocr
|
| 21 |
---
|
| 22 |
|
| 23 |
-
#
|
| 24 |
|
| 25 |
-
|
| 26 |
-
(multilingual document parsing for English and 22 Indian languages).
|
| 27 |
-
Same two-stage layout as upstream:
|
| 28 |
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
|
| 34 |
```python
|
| 35 |
-
from mlx_vlm.models.indic_ocr
|
| 36 |
|
| 37 |
with IndicOCRParser.from_pretrained("HashNuke/indic-ocr-mlx") as parser:
|
| 38 |
page = parser.parse("page.png")
|
|
@@ -41,16 +74,55 @@ print(page.markdown)
|
|
| 41 |
page.save("page.json")
|
| 42 |
```
|
| 43 |
|
| 44 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
|
| 46 |
```python
|
| 47 |
from pathlib import Path
|
|
|
|
|
|
|
| 48 |
from mlx_vlm import load
|
| 49 |
from mlx_vlm.utils import load_model
|
| 50 |
|
| 51 |
-
|
| 52 |
-
|
|
|
|
|
|
|
| 53 |
```
|
| 54 |
|
| 55 |
-
|
| 56 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
- en
|
| 4 |
- as
|
| 5 |
- bn
|
| 6 |
+
- brx
|
| 7 |
+
- doi
|
| 8 |
+
- gu
|
| 9 |
- hi
|
| 10 |
+
- kn
|
| 11 |
+
- ks
|
| 12 |
+
- kok
|
| 13 |
+
- mai
|
| 14 |
+
- ml
|
| 15 |
+
- mni
|
| 16 |
- mr
|
| 17 |
+
- ne
|
| 18 |
+
- or
|
| 19 |
+
- pa
|
| 20 |
+
- sa
|
| 21 |
+
- sat
|
| 22 |
+
- sd
|
| 23 |
- ta
|
| 24 |
- te
|
| 25 |
+
- ur
|
| 26 |
tags:
|
| 27 |
- mlx
|
| 28 |
- ocr
|
| 29 |
- document-parsing
|
| 30 |
+
- layout-analysis
|
| 31 |
+
- reading-order
|
| 32 |
- indic
|
| 33 |
library_name: mlx-vlm
|
| 34 |
pipeline_tag: image-text-to-text
|
| 35 |
license: other
|
| 36 |
license_name: indic-open-model-license-1.0
|
| 37 |
+
license_link: LICENSE.md
|
| 38 |
base_model: bodhan-ai/indic-ocr
|
| 39 |
---
|
| 40 |
|
| 41 |
+
# IndicOCR (MLX)
|
| 42 |
|
| 43 |
+
Built with IndicOCR from Bodhan AI / AI4Bharat.
|
|
|
|
|
|
|
| 44 |
|
| 45 |
+
MLX conversion of [bodhan-ai/indic-ocr](https://huggingface.co/bodhan-ai/indic-ocr)
|
| 46 |
+
for document parsing on Apple Silicon. A page image becomes reading-ordered
|
| 47 |
+
Markdown and per-block JSON, with equations in LaTeX and tables in HTML by default.
|
| 48 |
+
|
| 49 |
+
This repository contains two models:
|
| 50 |
+
|
| 51 |
+
| Stage | Directory | Model | Weights |
|
| 52 |
+
|---|---|---|---|
|
| 53 |
+
| Layout detection and reading order | `weights/layout` | IndicDocLayout: PP-DocLayoutV3 fine-tune, 37 classes, about 33M parameters | float32, 133 MB |
|
| 54 |
+
| Block transcription | `weights/ocr` | IndicBlockOCR: Qwen3.5-0.8B fine-tune | BF16, 1.7 GB, not quantized |
|
| 55 |
+
|
| 56 |
+
## Usage
|
| 57 |
+
|
| 58 |
+
Use an [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) checkout that includes
|
| 59 |
+
`indic_ocr` and `pp_doclayout_v3` support. From that checkout:
|
| 60 |
+
|
| 61 |
+
```sh
|
| 62 |
+
python -m pip install -e .
|
| 63 |
+
```
|
| 64 |
+
|
| 65 |
+
Replace `page.png` with a document image:
|
| 66 |
|
| 67 |
```python
|
| 68 |
+
from mlx_vlm.models.indic_ocr import IndicOCRParser
|
| 69 |
|
| 70 |
with IndicOCRParser.from_pretrained("HashNuke/indic-ocr-mlx") as parser:
|
| 71 |
page = parser.parse("page.png")
|
|
|
|
| 74 |
page.save("page.json")
|
| 75 |
```
|
| 76 |
|
| 77 |
+
The pipeline detects and cleans the layout, resolves nested equations, and
|
| 78 |
+
transcribes eligible blocks using greedy decoding. Skipped regions remain in
|
| 79 |
+
the JSON with empty text. Closing the parser releases its model references.
|
| 80 |
+
|
| 81 |
+
`page.save` writes the image name, dimensions, and block records; Markdown is
|
| 82 |
+
available separately as `page.markdown`. Each block has a zero-based `order`,
|
| 83 |
+
`label`, `type`, pixel-coordinate `bbox_xyxy` (`[x0, y0, x1, y1]`), confidence,
|
| 84 |
+
and text.
|
| 85 |
+
|
| 86 |
+
## Load individual stages
|
| 87 |
+
|
| 88 |
+
The repository root is a two-stage wrapper, not a standalone OCR model.
|
| 89 |
+
Download the repository and pass its local stage directories to the loaders:
|
| 90 |
|
| 91 |
```python
|
| 92 |
from pathlib import Path
|
| 93 |
+
|
| 94 |
+
from huggingface_hub import snapshot_download
|
| 95 |
from mlx_vlm import load
|
| 96 |
from mlx_vlm.utils import load_model
|
| 97 |
|
| 98 |
+
root = Path(snapshot_download("HashNuke/indic-ocr-mlx"))
|
| 99 |
+
layout = load_model(root / "weights/layout")
|
| 100 |
+
layout.eval()
|
| 101 |
+
ocr, processor = load(str(root / "weights/ocr"))
|
| 102 |
```
|
| 103 |
|
| 104 |
+
Do not pass `HashNuke/indic-ocr-mlx/weights/ocr` as a repository ID.
|
| 105 |
+
For stage-specific examples, see the
|
| 106 |
+
[layout README](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/weights/layout/README.md)
|
| 107 |
+
and [OCR README](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/weights/ocr/README.md).
|
| 108 |
+
|
| 109 |
+
## Languages and validation
|
| 110 |
+
|
| 111 |
+
Upstream reports printed-text support for English and 22 Indian languages.
|
| 112 |
+
Handwriting support covers English and 12 Indian languages; it does not cover
|
| 113 |
+
every printed-text language. See the
|
| 114 |
+
[upstream language coverage](https://huggingface.co/bodhan-ai/indic-ocr#supported-languages)
|
| 115 |
+
for details.
|
| 116 |
+
|
| 117 |
+
Greedy BF16 OCR matched fresh PyTorch transcriptions on an English title, an
|
| 118 |
+
equation, and a Telugu line. Complete page parsing was checked on an English
|
| 119 |
+
paper and annotated Telugu/Hindi gallery panels. A controlled table produced
|
| 120 |
+
HTML with the expected rows and cell values. These are sample-level checks,
|
| 121 |
+
not accuracy measurements across all supported languages; the gallery panels
|
| 122 |
+
retain annotation text and handwriting quality varies.
|
| 123 |
+
|
| 124 |
+
## License
|
| 125 |
+
|
| 126 |
+
The source weights are distributed under the
|
| 127 |
+
[Indic Open Model License v1.0](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/LICENSE.md).
|
| 128 |
+
This conversion does not change the upstream license.
|
weights/layout/README.md
CHANGED
|
@@ -1,26 +1,48 @@
|
|
| 1 |
---
|
| 2 |
-
language:
|
| 3 |
-
- en
|
| 4 |
tags:
|
| 5 |
- mlx
|
| 6 |
- object-detection
|
| 7 |
- document-layout
|
|
|
|
| 8 |
library_name: mlx-vlm
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
---
|
| 10 |
|
| 11 |
-
#
|
| 12 |
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
|
| 17 |
```python
|
| 18 |
from pathlib import Path
|
|
|
|
|
|
|
| 19 |
from mlx_vlm.utils import load_model
|
| 20 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 21 |
model.eval()
|
| 22 |
-
|
|
|
|
|
|
|
| 23 |
```
|
| 24 |
|
| 25 |
-
|
| 26 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
|
|
|
|
|
|
| 2 |
tags:
|
| 3 |
- mlx
|
| 4 |
- object-detection
|
| 5 |
- document-layout
|
| 6 |
+
- reading-order
|
| 7 |
library_name: mlx-vlm
|
| 8 |
+
license: other
|
| 9 |
+
license_name: indic-open-model-license-1.0
|
| 10 |
+
license_link: https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/LICENSE.md
|
| 11 |
+
base_model: bodhan-ai/indic-ocr
|
| 12 |
---
|
| 13 |
|
| 14 |
+
# IndicDocLayout (MLX)
|
| 15 |
|
| 16 |
+
Built with IndicDocLayout from Bodhan AI / AI4Bharat.
|
| 17 |
+
|
| 18 |
+
The layout stage of [IndicOCR (MLX)](https://huggingface.co/HashNuke/indic-ocr-mlx):
|
| 19 |
+
a 37-class PP-DocLayoutV3 fine-tune that predicts document regions and reading
|
| 20 |
+
order. The weights are float32, about 33M parameters (133 MB).
|
| 21 |
+
|
| 22 |
+
Use an mlx-vlm checkout with `pp_doclayout_v3` support. This example downloads
|
| 23 |
+
only the layout stage; it does not load the OCR model.
|
| 24 |
|
| 25 |
```python
|
| 26 |
from pathlib import Path
|
| 27 |
+
|
| 28 |
+
from huggingface_hub import snapshot_download
|
| 29 |
from mlx_vlm.utils import load_model
|
| 30 |
+
|
| 31 |
+
root = Path(snapshot_download(
|
| 32 |
+
"HashNuke/indic-ocr-mlx", allow_patterns=["weights/layout/*"]
|
| 33 |
+
))
|
| 34 |
+
model = load_model(root / "weights/layout")
|
| 35 |
model.eval()
|
| 36 |
+
records = model.detect("page.png", conf=0.5)
|
| 37 |
+
for record in sorted(records, key=lambda item: item["reading_order"]):
|
| 38 |
+
print(record)
|
| 39 |
```
|
| 40 |
|
| 41 |
+
Replace `page.png` with an image path. Records contain `bbox` in
|
| 42 |
+
`[y0, x0, y1, x1]` order, normalized to 0–1000, plus `label`, one-based
|
| 43 |
+
`reading_order`, and `score`. The detector does not transcribe text.
|
| 44 |
+
|
| 45 |
+
This fine-tune is distinct from the stock 25-class
|
| 46 |
+
[PP-DocLayout V3 (MLX)](https://huggingface.co/HashNuke/pp-doclayout-v3-mlx).
|
| 47 |
+
The source weights are from [bodhan-ai/indic-ocr](https://huggingface.co/bodhan-ai/indic-ocr)
|
| 48 |
+
under the [Indic Open Model License v1.0](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/LICENSE.md).
|
weights/ocr/README.md
CHANGED
|
@@ -1,41 +1,59 @@
|
|
| 1 |
---
|
| 2 |
-
language:
|
| 3 |
-
- en
|
| 4 |
-
- as
|
| 5 |
-
- bn
|
| 6 |
-
- hi
|
| 7 |
-
- mr
|
| 8 |
-
- ta
|
| 9 |
-
- te
|
| 10 |
tags:
|
| 11 |
- mlx
|
| 12 |
- ocr
|
| 13 |
- indic
|
| 14 |
library_name: mlx-vlm
|
| 15 |
pipeline_tag: image-text-to-text
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
---
|
| 17 |
|
| 18 |
-
#
|
| 19 |
|
| 20 |
-
|
| 21 |
-
[bodhan-ai/indic-ocr](https://huggingface.co/bodhan-ai/indic-ocr)
|
| 22 |
-
(`weights/ocr`, Qwen3.5-0.8B, bf16, **not quantized**).
|
| 23 |
|
| 24 |
-
|
| 25 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 26 |
|
| 27 |
```python
|
| 28 |
-
from mlx_vlm import load
|
| 29 |
-
from mlx_vlm.models.indic_ocr.pipeline import IndicOCRParser
|
| 30 |
-
from mlx_vlm.utils import load_model
|
| 31 |
from pathlib import Path
|
| 32 |
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
```
|
| 38 |
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
tags:
|
| 3 |
- mlx
|
| 4 |
- ocr
|
| 5 |
- indic
|
| 6 |
library_name: mlx-vlm
|
| 7 |
pipeline_tag: image-text-to-text
|
| 8 |
+
license: other
|
| 9 |
+
license_name: indic-open-model-license-1.0
|
| 10 |
+
license_link: https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/LICENSE.md
|
| 11 |
+
base_model: bodhan-ai/indic-ocr
|
| 12 |
---
|
| 13 |
|
| 14 |
+
# IndicBlockOCR (MLX)
|
| 15 |
|
| 16 |
+
Built with IndicBlockOCR from Bodhan AI / AI4Bharat.
|
|
|
|
|
|
|
| 17 |
|
| 18 |
+
The recognition stage of [IndicOCR (MLX)](https://huggingface.co/HashNuke/indic-ocr-mlx):
|
| 19 |
+
a Qwen3.5-0.8B fine-tune for transcribing text, equations, and tables from
|
| 20 |
+
cropped document regions. The weights are BF16 (1.7 GB), not quantized.
|
| 21 |
+
|
| 22 |
+
Use an mlx-vlm checkout with `indic_ocr` support. Replace `crop.png` with an
|
| 23 |
+
image of a text region, not a complete multi-region page:
|
| 24 |
|
| 25 |
```python
|
|
|
|
|
|
|
|
|
|
| 26 |
from pathlib import Path
|
| 27 |
|
| 28 |
+
from huggingface_hub import snapshot_download
|
| 29 |
+
from PIL import Image
|
| 30 |
+
from mlx_vlm import generate, load
|
| 31 |
+
from mlx_vlm.models.indic_ocr.processing_indic_ocr import area_clamp, prompt_for
|
| 32 |
+
from mlx_vlm.prompt_utils import apply_chat_template
|
| 33 |
+
|
| 34 |
+
root = Path(snapshot_download(
|
| 35 |
+
"HashNuke/indic-ocr-mlx", allow_patterns=["weights/ocr/*"]
|
| 36 |
+
))
|
| 37 |
+
model, processor = load(str(root / "weights/ocr"))
|
| 38 |
+
with Image.open("crop.png") as image:
|
| 39 |
+
crop = area_clamp(image.convert("RGB"))
|
| 40 |
+
prompt = apply_chat_template(
|
| 41 |
+
processor, model.config, prompt_for("Text"), num_images=1
|
| 42 |
+
)
|
| 43 |
+
result = generate(
|
| 44 |
+
model, processor, prompt, [crop],
|
| 45 |
+
max_tokens=2048, temperature=0.0, verbose=False,
|
| 46 |
+
)
|
| 47 |
+
print(result.text)
|
| 48 |
```
|
| 49 |
|
| 50 |
+
Use `prompt_for("Equation")` for LaTeX or `prompt_for("Table")` for HTML
|
| 51 |
+
tables. Greedy decoding (`temperature=0.0`) matches the default page pipeline.
|
| 52 |
+
|
| 53 |
+
The repository root is a two-stage wrapper: `load("HashNuke/indic-ocr-mlx")`
|
| 54 |
+
does not load this OCR stage. Use the local `weights/ocr` path above, or
|
| 55 |
+
`IndicOCRParser.from_pretrained("HashNuke/indic-ocr-mlx")` for complete pages,
|
| 56 |
+
as shown in the [main model card](https://huggingface.co/HashNuke/indic-ocr-mlx).
|
| 57 |
+
|
| 58 |
+
The source weights are from [bodhan-ai/indic-ocr](https://huggingface.co/bodhan-ai/indic-ocr)
|
| 59 |
+
under the [Indic Open Model License v1.0](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/LICENSE.md).
|