Object Detection
Safetensors
MLX
English
Chinese
mlx-vlm
pp_doclayout_v3
document-layout
reading-order
Instructions to use HashNuke/pp-doclayout-v3-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use HashNuke/pp-doclayout-v3-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir pp-doclayout-v3-mlx HashNuke/pp-doclayout-v3-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
|
Download README.md from HashNuke/pp-doclayout-v3-mlx: direct link, hf CLI and curl.
- Browser
- Download file 2.91 kB
-
https://huggingface.co/HashNuke/pp-doclayout-v3-mlx/resolve/main/README.md
- Command line
-
hf download hf://HashNuke/pp-doclayout-v3-mlx/README.md
-
curl -L -o README.md https://huggingface.co/HashNuke/pp-doclayout-v3-mlx/resolve/main/README.md
2.91 kB
| language: | |
| - en | |
| - zh | |
| tags: | |
| - mlx | |
| - object-detection | |
| - document-layout | |
| - reading-order | |
| library_name: mlx-vlm | |
| pipeline_tag: object-detection | |
| license: apache-2.0 | |
| base_model: PaddlePaddle/PP-DocLayoutV3_safetensors | |
| # PP-DocLayout V3 (MLX) | |
| MLX conversion of [PaddlePaddle/PP-DocLayoutV3_safetensors](https://huggingface.co/PaddlePaddle/PP-DocLayoutV3_safetensors) | |
| for document-layout detection on Apple Silicon. It predicts page regions, | |
| class labels, and reading order; it does not transcribe text. | |
| - **Architecture:** HGNetV2-L backbone, hybrid encoder, and deformable decoder | |
| with a reading-order head. | |
| - **Classes:** 25 prediction classes from the stock checkpoint. | |
| - **Weights:** about 33M parameters, float32, 133 MB, not quantized. | |
| ## Usage | |
| Use an [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) checkout that includes | |
| `pp_doclayout_v3` support. From that checkout: | |
| ```sh | |
| python -m pip install -e . | |
| ``` | |
| Replace `page.png` with a document image: | |
| ```python | |
| from mlx_vlm.utils import get_model_path, load_model | |
| model = load_model(get_model_path("HashNuke/pp-doclayout-v3-mlx")) | |
| model.eval() | |
| records = model.detect("page.png", conf=0.5) | |
| for record in sorted(records, key=lambda item: item["reading_order"]): | |
| print(record) | |
| ``` | |
| `get_model_path` downloads the checkpoint and returns a local path for | |
| `load_model`. `detect` accepts an image path or a PIL image. | |
| ## Output format | |
| Each detection contains: | |
| - `bbox`: `[y0, x0, y1, x1]`, normalized to 0–1000 and rounded to one decimal. | |
| - `label`: the checkpoint's class name. | |
| - `reading_order`: a one-based rank among retained detections. | |
| - `score`: confidence rounded to three decimals. | |
| The returned list is in query order; sort by `reading_order` to consume it in | |
| document order. This MLX interface returns rectangular boxes, not polygons or | |
| segmentation masks. | |
| The default inference recipe resizes RGB images to 1024×1024 and scales pixels | |
| to [0, 1]. This matches IndicDocLayout's preprocessing and differs from the | |
| stock Transformers processor's defaults. The mask-feature branch is retained | |
| for mask-enhanced query initialization; training-only denoising weights are omitted. | |
| ## Validation | |
| At `conf=0.5` with the same preprocessing and reading-order decoding, output | |
| records matched fresh PyTorch reference runs on the English paper and | |
| calendar/table samples in mlx-vlm (`examples/images/paper.png` and | |
| `examples/images/demo_pdf1_page1.png`). Labels, reading order, and rounded | |
| boxes/scores matched. This is sample-level output parity, not bitwise equality | |
| of raw tensors or a detection-accuracy benchmark. | |
| The 37-class IndicDocLayout fine-tune is available separately as the layout | |
| stage of [IndicOCR (MLX)](https://huggingface.co/HashNuke/indic-ocr-mlx). | |
| ## License | |
| The [upstream weights](https://huggingface.co/PaddlePaddle/PP-DocLayoutV3_safetensors) | |
| are licensed under Apache-2.0. This conversion does not change that license. | |