HashNuke commited on
Commit
3ca5518
·
verified ·
1 Parent(s): 6a0d82c

Clarify IndicOCR usage and stage loading

Browse files

Document page parsing and standalone stage inference. Correct model metadata, license links, and validation scope.

Files changed (3) hide show
  1. README.md +87 -15
  2. weights/layout/README.md +32 -10
  3. weights/ocr/README.md +42 -24
README.md CHANGED
@@ -3,36 +3,69 @@ language:
3
  - en
4
  - as
5
  - bn
 
 
 
6
  - hi
 
 
 
 
 
 
7
  - mr
 
 
 
 
 
 
8
  - ta
9
  - te
 
10
  tags:
11
  - mlx
12
  - ocr
13
  - document-parsing
 
 
14
  - indic
15
  library_name: mlx-vlm
16
  pipeline_tag: image-text-to-text
17
  license: other
18
  license_name: indic-open-model-license-1.0
19
- license_link: LICENSE
20
  base_model: bodhan-ai/indic-ocr
21
  ---
22
 
23
- # indic-ocr-mlx
24
 
25
- MLX port of [bodhan-ai/indic-ocr](https://huggingface.co/bodhan-ai/indic-ocr)
26
- (multilingual document parsing for English and 22 Indian languages).
27
- Same two-stage layout as upstream:
28
 
29
- | Stage | Dir | Model | Format |
30
- |-------|-----|-------|--------|
31
- | IndicDocLayout | `weights/layout` | PP-DocLayoutV3/RT-DETR, 37 classes + reading order | float32 |
32
- | IndicBlockOCR | `weights/ocr` | Qwen3.5-0.8B | bf16, not quantized |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
33
 
34
  ```python
35
- from mlx_vlm.models.indic_ocr.pipeline import IndicOCRParser
36
 
37
  with IndicOCRParser.from_pretrained("HashNuke/indic-ocr-mlx") as parser:
38
  page = parser.parse("page.png")
@@ -41,16 +74,55 @@ print(page.markdown)
41
  page.save("page.json")
42
  ```
43
 
44
- Stages also load standalone with standard tooling:
 
 
 
 
 
 
 
 
 
 
 
 
45
 
46
  ```python
47
  from pathlib import Path
 
 
48
  from mlx_vlm import load
49
  from mlx_vlm.utils import load_model
50
 
51
- layout = load_model(Path("HashNuke/indic-ocr-mlx/weights/layout"))
52
- ocr, processor = load("HashNuke/indic-ocr-mlx/weights/ocr")
 
 
53
  ```
54
 
55
- Original weights: `bodhan-ai/indic-ocr` (gated, Indic Open Model License
56
- v1.0). Port code: `mlx_vlm/models/indic_ocr` in mlx-vlm.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  - en
4
  - as
5
  - bn
6
+ - brx
7
+ - doi
8
+ - gu
9
  - hi
10
+ - kn
11
+ - ks
12
+ - kok
13
+ - mai
14
+ - ml
15
+ - mni
16
  - mr
17
+ - ne
18
+ - or
19
+ - pa
20
+ - sa
21
+ - sat
22
+ - sd
23
  - ta
24
  - te
25
+ - ur
26
  tags:
27
  - mlx
28
  - ocr
29
  - document-parsing
30
+ - layout-analysis
31
+ - reading-order
32
  - indic
33
  library_name: mlx-vlm
34
  pipeline_tag: image-text-to-text
35
  license: other
36
  license_name: indic-open-model-license-1.0
37
+ license_link: LICENSE.md
38
  base_model: bodhan-ai/indic-ocr
39
  ---
40
 
41
+ # IndicOCR (MLX)
42
 
43
+ Built with IndicOCR from Bodhan AI / AI4Bharat.
 
 
44
 
45
+ MLX conversion of [bodhan-ai/indic-ocr](https://huggingface.co/bodhan-ai/indic-ocr)
46
+ for document parsing on Apple Silicon. A page image becomes reading-ordered
47
+ Markdown and per-block JSON, with equations in LaTeX and tables in HTML by default.
48
+
49
+ This repository contains two models:
50
+
51
+ | Stage | Directory | Model | Weights |
52
+ |---|---|---|---|
53
+ | Layout detection and reading order | `weights/layout` | IndicDocLayout: PP-DocLayoutV3 fine-tune, 37 classes, about 33M parameters | float32, 133 MB |
54
+ | Block transcription | `weights/ocr` | IndicBlockOCR: Qwen3.5-0.8B fine-tune | BF16, 1.7 GB, not quantized |
55
+
56
+ ## Usage
57
+
58
+ Use an [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) checkout that includes
59
+ `indic_ocr` and `pp_doclayout_v3` support. From that checkout:
60
+
61
+ ```sh
62
+ python -m pip install -e .
63
+ ```
64
+
65
+ Replace `page.png` with a document image:
66
 
67
  ```python
68
+ from mlx_vlm.models.indic_ocr import IndicOCRParser
69
 
70
  with IndicOCRParser.from_pretrained("HashNuke/indic-ocr-mlx") as parser:
71
  page = parser.parse("page.png")
 
74
  page.save("page.json")
75
  ```
76
 
77
+ The pipeline detects and cleans the layout, resolves nested equations, and
78
+ transcribes eligible blocks using greedy decoding. Skipped regions remain in
79
+ the JSON with empty text. Closing the parser releases its model references.
80
+
81
+ `page.save` writes the image name, dimensions, and block records; Markdown is
82
+ available separately as `page.markdown`. Each block has a zero-based `order`,
83
+ `label`, `type`, pixel-coordinate `bbox_xyxy` (`[x0, y0, x1, y1]`), confidence,
84
+ and text.
85
+
86
+ ## Load individual stages
87
+
88
+ The repository root is a two-stage wrapper, not a standalone OCR model.
89
+ Download the repository and pass its local stage directories to the loaders:
90
 
91
  ```python
92
  from pathlib import Path
93
+
94
+ from huggingface_hub import snapshot_download
95
  from mlx_vlm import load
96
  from mlx_vlm.utils import load_model
97
 
98
+ root = Path(snapshot_download("HashNuke/indic-ocr-mlx"))
99
+ layout = load_model(root / "weights/layout")
100
+ layout.eval()
101
+ ocr, processor = load(str(root / "weights/ocr"))
102
  ```
103
 
104
+ Do not pass `HashNuke/indic-ocr-mlx/weights/ocr` as a repository ID.
105
+ For stage-specific examples, see the
106
+ [layout README](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/weights/layout/README.md)
107
+ and [OCR README](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/weights/ocr/README.md).
108
+
109
+ ## Languages and validation
110
+
111
+ Upstream reports printed-text support for English and 22 Indian languages.
112
+ Handwriting support covers English and 12 Indian languages; it does not cover
113
+ every printed-text language. See the
114
+ [upstream language coverage](https://huggingface.co/bodhan-ai/indic-ocr#supported-languages)
115
+ for details.
116
+
117
+ Greedy BF16 OCR matched fresh PyTorch transcriptions on an English title, an
118
+ equation, and a Telugu line. Complete page parsing was checked on an English
119
+ paper and annotated Telugu/Hindi gallery panels. A controlled table produced
120
+ HTML with the expected rows and cell values. These are sample-level checks,
121
+ not accuracy measurements across all supported languages; the gallery panels
122
+ retain annotation text and handwriting quality varies.
123
+
124
+ ## License
125
+
126
+ The source weights are distributed under the
127
+ [Indic Open Model License v1.0](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/LICENSE.md).
128
+ This conversion does not change the upstream license.
weights/layout/README.md CHANGED
@@ -1,26 +1,48 @@
1
  ---
2
- language:
3
- - en
4
  tags:
5
  - mlx
6
  - object-detection
7
  - document-layout
 
8
  library_name: mlx-vlm
 
 
 
 
9
  ---
10
 
11
- # indic-layout-mlx (layout stage, float32)
12
 
13
- MLX conversion of the **IndicDocLayout** stage of
14
- [bodhan-ai/indic-ocr](https://huggingface.co/bodhan-ai/indic-ocr)
15
- (`weights/layout`, PP-DocLayoutV3/RT-DETR 33M, 37 classes + reading order).
 
 
 
 
 
16
 
17
  ```python
18
  from pathlib import Path
 
 
19
  from mlx_vlm.utils import load_model
20
- model = load_model(Path("HashNuke/indic-layout-mlx"))
 
 
 
 
21
  model.eval()
22
- print(model.detect("page.png", conf=0.5))
 
 
23
  ```
24
 
25
- Original model: `bodhan-ai/indic-ocr` (gated, Indic Open Model License v1.0).
26
- Converted with `mlx_vlm.models.indic_ocr.convert_layout --dtype float32`.
 
 
 
 
 
 
 
1
  ---
 
 
2
  tags:
3
  - mlx
4
  - object-detection
5
  - document-layout
6
+ - reading-order
7
  library_name: mlx-vlm
8
+ license: other
9
+ license_name: indic-open-model-license-1.0
10
+ license_link: https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/LICENSE.md
11
+ base_model: bodhan-ai/indic-ocr
12
  ---
13
 
14
+ # IndicDocLayout (MLX)
15
 
16
+ Built with IndicDocLayout from Bodhan AI / AI4Bharat.
17
+
18
+ The layout stage of [IndicOCR (MLX)](https://huggingface.co/HashNuke/indic-ocr-mlx):
19
+ a 37-class PP-DocLayoutV3 fine-tune that predicts document regions and reading
20
+ order. The weights are float32, about 33M parameters (133 MB).
21
+
22
+ Use an mlx-vlm checkout with `pp_doclayout_v3` support. This example downloads
23
+ only the layout stage; it does not load the OCR model.
24
 
25
  ```python
26
  from pathlib import Path
27
+
28
+ from huggingface_hub import snapshot_download
29
  from mlx_vlm.utils import load_model
30
+
31
+ root = Path(snapshot_download(
32
+ "HashNuke/indic-ocr-mlx", allow_patterns=["weights/layout/*"]
33
+ ))
34
+ model = load_model(root / "weights/layout")
35
  model.eval()
36
+ records = model.detect("page.png", conf=0.5)
37
+ for record in sorted(records, key=lambda item: item["reading_order"]):
38
+ print(record)
39
  ```
40
 
41
+ Replace `page.png` with an image path. Records contain `bbox` in
42
+ `[y0, x0, y1, x1]` order, normalized to 0–1000, plus `label`, one-based
43
+ `reading_order`, and `score`. The detector does not transcribe text.
44
+
45
+ This fine-tune is distinct from the stock 25-class
46
+ [PP-DocLayout V3 (MLX)](https://huggingface.co/HashNuke/pp-doclayout-v3-mlx).
47
+ The source weights are from [bodhan-ai/indic-ocr](https://huggingface.co/bodhan-ai/indic-ocr)
48
+ under the [Indic Open Model License v1.0](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/LICENSE.md).
weights/ocr/README.md CHANGED
@@ -1,41 +1,59 @@
1
  ---
2
- language:
3
- - en
4
- - as
5
- - bn
6
- - hi
7
- - mr
8
- - ta
9
- - te
10
  tags:
11
  - mlx
12
  - ocr
13
  - indic
14
  library_name: mlx-vlm
15
  pipeline_tag: image-text-to-text
 
 
 
 
16
  ---
17
 
18
- # indic-ocr-mlx (OCR stage, bf16)
19
 
20
- MLX conversion of the **IndicBlockOCR** stage of
21
- [bodhan-ai/indic-ocr](https://huggingface.co/bodhan-ai/indic-ocr)
22
- (`weights/ocr`, Qwen3.5-0.8B, bf16, **not quantized**).
23
 
24
- Use with `mlx-vlm` `indic_ocr` model support
25
- (`mlx_vlm/models/indic_ocr`):
 
 
 
 
26
 
27
  ```python
28
- from mlx_vlm import load
29
- from mlx_vlm.models.indic_ocr.pipeline import IndicOCRParser
30
- from mlx_vlm.utils import load_model
31
  from pathlib import Path
32
 
33
- layout = load_model(Path("HashNuke/indic-layout-mlx"))
34
- ocr, processor = load("HashNuke/indic-ocr-mlx")
35
- page = IndicOCRParser(layout, ocr, processor).parse("page.png")
36
- print(page.markdown)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
37
  ```
38
 
39
- Original model: `bodhan-ai/indic-ocr` (gated, Indic Open Model License v1.0).
40
- Converted with `mlx_vlm.convert --dtype bfloat16`; `config.json`
41
- `model_type` rewritten to `indic_ocr`.
 
 
 
 
 
 
 
 
1
  ---
 
 
 
 
 
 
 
 
2
  tags:
3
  - mlx
4
  - ocr
5
  - indic
6
  library_name: mlx-vlm
7
  pipeline_tag: image-text-to-text
8
+ license: other
9
+ license_name: indic-open-model-license-1.0
10
+ license_link: https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/LICENSE.md
11
+ base_model: bodhan-ai/indic-ocr
12
  ---
13
 
14
+ # IndicBlockOCR (MLX)
15
 
16
+ Built with IndicBlockOCR from Bodhan AI / AI4Bharat.
 
 
17
 
18
+ The recognition stage of [IndicOCR (MLX)](https://huggingface.co/HashNuke/indic-ocr-mlx):
19
+ a Qwen3.5-0.8B fine-tune for transcribing text, equations, and tables from
20
+ cropped document regions. The weights are BF16 (1.7 GB), not quantized.
21
+
22
+ Use an mlx-vlm checkout with `indic_ocr` support. Replace `crop.png` with an
23
+ image of a text region, not a complete multi-region page:
24
 
25
  ```python
 
 
 
26
  from pathlib import Path
27
 
28
+ from huggingface_hub import snapshot_download
29
+ from PIL import Image
30
+ from mlx_vlm import generate, load
31
+ from mlx_vlm.models.indic_ocr.processing_indic_ocr import area_clamp, prompt_for
32
+ from mlx_vlm.prompt_utils import apply_chat_template
33
+
34
+ root = Path(snapshot_download(
35
+ "HashNuke/indic-ocr-mlx", allow_patterns=["weights/ocr/*"]
36
+ ))
37
+ model, processor = load(str(root / "weights/ocr"))
38
+ with Image.open("crop.png") as image:
39
+ crop = area_clamp(image.convert("RGB"))
40
+ prompt = apply_chat_template(
41
+ processor, model.config, prompt_for("Text"), num_images=1
42
+ )
43
+ result = generate(
44
+ model, processor, prompt, [crop],
45
+ max_tokens=2048, temperature=0.0, verbose=False,
46
+ )
47
+ print(result.text)
48
  ```
49
 
50
+ Use `prompt_for("Equation")` for LaTeX or `prompt_for("Table")` for HTML
51
+ tables. Greedy decoding (`temperature=0.0`) matches the default page pipeline.
52
+
53
+ The repository root is a two-stage wrapper: `load("HashNuke/indic-ocr-mlx")`
54
+ does not load this OCR stage. Use the local `weights/ocr` path above, or
55
+ `IndicOCRParser.from_pretrained("HashNuke/indic-ocr-mlx")` for complete pages,
56
+ as shown in the [main model card](https://huggingface.co/HashNuke/indic-ocr-mlx).
57
+
58
+ The source weights are from [bodhan-ai/indic-ocr](https://huggingface.co/bodhan-ai/indic-ocr)
59
+ under the [Indic Open Model License v1.0](https://huggingface.co/HashNuke/indic-ocr-mlx/blob/main/LICENSE.md).