Object Detection
Safetensors
MLX
English
Chinese
mlx-vlm
pp_doclayout_v3
document-layout
reading-order
Instructions to use HashNuke/pp-doclayout-v3-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use HashNuke/pp-doclayout-v3-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir pp-doclayout-v3-mlx HashNuke/pp-doclayout-v3-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Clarify PP-DocLayout V3 usage
Browse filesDocument native MLX loading, output coordinates, and validation scope.
README.md
CHANGED
|
@@ -6,27 +6,79 @@ tags:
|
|
| 6 |
- mlx
|
| 7 |
- object-detection
|
| 8 |
- document-layout
|
|
|
|
| 9 |
library_name: mlx-vlm
|
|
|
|
|
|
|
|
|
|
| 10 |
---
|
| 11 |
|
| 12 |
-
#
|
| 13 |
|
| 14 |
-
MLX
|
| 15 |
-
|
| 16 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
|
| 18 |
```python
|
| 19 |
-
from mlx_vlm.utils import load_model
|
| 20 |
|
| 21 |
-
model = load_model("HashNuke/pp-doclayout-v3-mlx")
|
| 22 |
model.eval()
|
| 23 |
-
|
|
|
|
|
|
|
| 24 |
```
|
| 25 |
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 29 |
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
`python -m mlx_vlm.models.pp_doclayout_v3.convert --dtype float32`.
|
|
|
|
| 6 |
- mlx
|
| 7 |
- object-detection
|
| 8 |
- document-layout
|
| 9 |
+
- reading-order
|
| 10 |
library_name: mlx-vlm
|
| 11 |
+
pipeline_tag: object-detection
|
| 12 |
+
license: apache-2.0
|
| 13 |
+
base_model: PaddlePaddle/PP-DocLayoutV3_safetensors
|
| 14 |
---
|
| 15 |
|
| 16 |
+
# PP-DocLayout V3 (MLX)
|
| 17 |
|
| 18 |
+
MLX conversion of [PaddlePaddle/PP-DocLayoutV3_safetensors](https://huggingface.co/PaddlePaddle/PP-DocLayoutV3_safetensors)
|
| 19 |
+
for document-layout detection on Apple Silicon. It predicts page regions,
|
| 20 |
+
class labels, and reading order; it does not transcribe text.
|
| 21 |
+
|
| 22 |
+
- **Architecture:** HGNetV2-L backbone, hybrid encoder, and deformable decoder
|
| 23 |
+
with a reading-order head.
|
| 24 |
+
- **Classes:** 25 prediction classes from the stock checkpoint.
|
| 25 |
+
- **Weights:** about 33M parameters, float32, 133 MB, not quantized.
|
| 26 |
+
|
| 27 |
+
## Usage
|
| 28 |
+
|
| 29 |
+
Use an [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) checkout that includes
|
| 30 |
+
`pp_doclayout_v3` support. From that checkout:
|
| 31 |
+
|
| 32 |
+
```sh
|
| 33 |
+
python -m pip install -e .
|
| 34 |
+
```
|
| 35 |
+
|
| 36 |
+
Replace `page.png` with a document image:
|
| 37 |
|
| 38 |
```python
|
| 39 |
+
from mlx_vlm.utils import get_model_path, load_model
|
| 40 |
|
| 41 |
+
model = load_model(get_model_path("HashNuke/pp-doclayout-v3-mlx"))
|
| 42 |
model.eval()
|
| 43 |
+
records = model.detect("page.png", conf=0.5)
|
| 44 |
+
for record in sorted(records, key=lambda item: item["reading_order"]):
|
| 45 |
+
print(record)
|
| 46 |
```
|
| 47 |
|
| 48 |
+
`get_model_path` downloads the checkpoint and returns a local path for
|
| 49 |
+
`load_model`. `detect` accepts an image path or a PIL image.
|
| 50 |
+
|
| 51 |
+
## Output format
|
| 52 |
+
|
| 53 |
+
Each detection contains:
|
| 54 |
+
|
| 55 |
+
- `bbox`: `[y0, x0, y1, x1]`, normalized to 0–1000 and rounded to one decimal.
|
| 56 |
+
- `label`: the checkpoint's class name.
|
| 57 |
+
- `reading_order`: a one-based rank among retained detections.
|
| 58 |
+
- `score`: confidence rounded to three decimals.
|
| 59 |
+
|
| 60 |
+
The returned list is in query order; sort by `reading_order` to consume it in
|
| 61 |
+
document order. This MLX interface returns rectangular boxes, not polygons or
|
| 62 |
+
segmentation masks.
|
| 63 |
+
|
| 64 |
+
The default inference recipe resizes RGB images to 1024×1024 and scales pixels
|
| 65 |
+
to [0, 1]. This matches IndicDocLayout's preprocessing and differs from the
|
| 66 |
+
stock Transformers processor's defaults. The mask-feature branch is retained
|
| 67 |
+
for mask-enhanced query initialization; training-only denoising weights are omitted.
|
| 68 |
+
|
| 69 |
+
## Validation
|
| 70 |
+
|
| 71 |
+
At `conf=0.5` with the same preprocessing and reading-order decoding, output
|
| 72 |
+
records matched fresh PyTorch reference runs on the English paper and
|
| 73 |
+
calendar/table samples in mlx-vlm (`examples/images/paper.png` and
|
| 74 |
+
`examples/images/demo_pdf1_page1.png`). Labels, reading order, and rounded
|
| 75 |
+
boxes/scores matched. This is sample-level output parity, not bitwise equality
|
| 76 |
+
of raw tensors or a detection-accuracy benchmark.
|
| 77 |
+
|
| 78 |
+
The 37-class IndicDocLayout fine-tune is available separately as the layout
|
| 79 |
+
stage of [IndicOCR (MLX)](https://huggingface.co/HashNuke/indic-ocr-mlx).
|
| 80 |
+
|
| 81 |
+
## License
|
| 82 |
|
| 83 |
+
The [upstream weights](https://huggingface.co/PaddlePaddle/PP-DocLayoutV3_safetensors)
|
| 84 |
+
are licensed under Apache-2.0. This conversion does not change that license.
|
|
|