Image-Text-to-Text
Safetensors
MLX
mlx-vlm
indic_ocr
ocr
document-parsing
layout-analysis
reading-order
indic
Instructions to use HashNuke/indic-ocr-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use HashNuke/indic-ocr-mlx with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("HashNuke/indic-ocr-mlx") config = load_config("HashNuke/indic-ocr-mlx") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
File size: 3,977 Bytes
b68daba cbcbaad 3ca5518 cbcbaad 3ca5518 cbcbaad 3ca5518 cbcbaad 3ca5518 cbcbaad 3ca5518 cbcbaad b68daba 3ca5518 cbcbaad b68daba cbcbaad 3ca5518 cbcbaad 3ca5518 cbcbaad 3ca5518 04cb39f 3ca5518 cbcbaad 04cb39f cbcbaad 04cb39f cbcbaad 04cb39f 3ca5518 cbcbaad 3ca5518 cbcbaad 3ca5518 cbcbaad 3ca5518 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 | ---
language:
- en
- as
- bn
- brx
- doi
- gu
- hi
- kn
- ks
- kok
- mai
- ml
- mni
- mr
- ne
- or
- pa
- sa
- sat
- sd
- ta
- te
- ur
tags:
- mlx
- ocr
- document-parsing
- layout-analysis
- reading-order
- indic
library_name: mlx-vlm
pipeline_tag: image-text-to-text
license: other
license_name: indic-open-model-license-1.0
license_link: LICENSE.md
base_model: bodhan-ai/indic-ocr
---
# IndicOCR (MLX)
Built with IndicOCR from Bodhan AI / AI4Bharat.
MLX conversion of [bodhan-ai/indic-ocr](https://huggingface.co/bodhan-ai/indic-ocr)
for document parsing on Apple Silicon. A page image becomes reading-ordered
Markdown and per-block JSON, with equations in LaTeX and tables in HTML by default.
This repository contains two models:
| Stage | Directory | Model | Weights |
|---|---|---|---|
| Layout detection and reading order | `weights/layout` | IndicDocLayout: PP-DocLayoutV3 fine-tune, 37 classes, about 33M parameters | float32, 133 MB |
| Block transcription | `weights/ocr` | IndicBlockOCR: Qwen3.5-0.8B fine-tune | BF16, 1.7 GB, not quantized |
## Usage
Use an [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) checkout that includes
`indic_ocr` support for the standard `load_model` interface. From that checkout:
```sh
python -m pip install -e .
```
Replace `page.png` with a document image:
```python
from mlx_vlm.utils import get_model_path, load_model
with load_model(get_model_path("HashNuke/indic-ocr-mlx")) as parser:
page = parser.parse("page.png")
print(page.markdown)
page.save("page.json")
```
The root config embeds both stage configurations, and a standard safetensors
index points to their existing weight files. No conversion or preparation step
is needed. The weight files are unchanged.
The pipeline detects and cleans the layout, resolves nested equations, and
transcribes eligible blocks using greedy decoding. Skipped regions remain in
the JSON with empty text. Closing the parser releases its model references.
`page.save` writes the image name, dimensions, and block records; Markdown is
available separately as `page.markdown`. Each block has a zero-based `order`,
`label`, `type`, pixel-coordinate `bbox_xyxy` (`[x0, y0, x1, y1]`), confidence,
and text.
## Load individual stages
The repository root is a two-stage wrapper, not a standalone OCR model.
Download the repository and pass its local stage directories to the loaders:
```python
from pathlib import Path
from huggingface_hub import snapshot_download
from mlx_vlm import load
from mlx_vlm.utils import load_model
root = Path(snapshot_download("HashNuke/indic-ocr-mlx"))
layout = load_model(root / "weights/layout")
layout.eval()
ocr, processor = load(str(root / "weights/ocr"))
```
Do not pass `HashNuke/indic-ocr-mlx/weights/ocr` as a repository ID.
For stage-specific examples, see the
[layout README](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/weights/layout/README.md)
and [OCR README](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/weights/ocr/README.md).
## Languages and validation
Upstream reports printed-text support for English and 22 Indian languages.
Handwriting support covers English and 12 Indian languages; it does not cover
every printed-text language. See the
[upstream language coverage](https://huggingface.co/bodhan-ai/indic-ocr#supported-languages)
for details.
Greedy BF16 OCR matched fresh PyTorch transcriptions on an English title, an
equation, and a Telugu line. Complete page parsing was checked on an English
paper and annotated Telugu/Hindi gallery panels. A controlled table produced
HTML with the expected rows and cell values. These are sample-level checks,
not accuracy measurements across all supported languages; the gallery panels
retain annotation text and handwriting quality varies.
## License
The source weights are distributed under the
[Indic Open Model License v1.0](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/LICENSE.md).
This conversion does not change the upstream license.
|