Image-Text-to-Text
Safetensors
MLX
mlx-vlm
indic_ocr
ocr
document-parsing
layout-analysis
reading-order
indic
Instructions to use HashNuke/indic-ocr-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use HashNuke/indic-ocr-mlx with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("HashNuke/indic-ocr-mlx") config = load_config("HashNuke/indic-ocr-mlx") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Akash Manohar commited on
Support standard mlx-vlm model loading
Browse filesEmbed the layout and OCR configurations and add a standard safetensors index for the existing stage files. Map tensor keys to their parser submodules and document direct load_model usage without a preparation step.
- README.md +7 -3
- config.json +0 -0
- model.safetensors.index.json +0 -0
README.md
CHANGED
|
@@ -56,7 +56,7 @@ This repository contains two models:
|
|
| 56 |
## Usage
|
| 57 |
|
| 58 |
Use an [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) checkout that includes
|
| 59 |
-
`indic_ocr`
|
| 60 |
|
| 61 |
```sh
|
| 62 |
python -m pip install -e .
|
|
@@ -65,15 +65,19 @@ python -m pip install -e .
|
|
| 65 |
Replace `page.png` with a document image:
|
| 66 |
|
| 67 |
```python
|
| 68 |
-
from mlx_vlm.
|
| 69 |
|
| 70 |
-
with
|
| 71 |
page = parser.parse("page.png")
|
| 72 |
|
| 73 |
print(page.markdown)
|
| 74 |
page.save("page.json")
|
| 75 |
```
|
| 76 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 77 |
The pipeline detects and cleans the layout, resolves nested equations, and
|
| 78 |
transcribes eligible blocks using greedy decoding. Skipped regions remain in
|
| 79 |
the JSON with empty text. Closing the parser releases its model references.
|
|
|
|
| 56 |
## Usage
|
| 57 |
|
| 58 |
Use an [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) checkout that includes
|
| 59 |
+
`indic_ocr` support for the standard `load_model` interface. From that checkout:
|
| 60 |
|
| 61 |
```sh
|
| 62 |
python -m pip install -e .
|
|
|
|
| 65 |
Replace `page.png` with a document image:
|
| 66 |
|
| 67 |
```python
|
| 68 |
+
from mlx_vlm.utils import get_model_path, load_model
|
| 69 |
|
| 70 |
+
with load_model(get_model_path("HashNuke/indic-ocr-mlx")) as parser:
|
| 71 |
page = parser.parse("page.png")
|
| 72 |
|
| 73 |
print(page.markdown)
|
| 74 |
page.save("page.json")
|
| 75 |
```
|
| 76 |
|
| 77 |
+
The root config embeds both stage configurations, and a standard safetensors
|
| 78 |
+
index points to their existing weight files. No conversion or preparation step
|
| 79 |
+
is needed. The weight files are unchanged.
|
| 80 |
+
|
| 81 |
The pipeline detects and cleans the layout, resolves nested equations, and
|
| 82 |
transcribes eligible blocks using greedy decoding. Skipped regions remain in
|
| 83 |
the JSON with empty text. Closing the parser releases its model references.
|
config.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
model.safetensors.index.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|