Object Detection
Safetensors
MLX
English
Chinese
mlx-vlm
pp_doclayout_v3
document-layout
reading-order
Instructions to use HashNuke/pp-doclayout-v3-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use HashNuke/pp-doclayout-v3-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir pp-doclayout-v3-mlx HashNuke/pp-doclayout-v3-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
File size: 2,913 Bytes
9a34dcb cd50996 9a34dcb cd50996 9a34dcb cd50996 9a34dcb cd50996 9a34dcb cd50996 9a34dcb cd50996 9a34dcb cd50996 9a34dcb cd50996 9a34dcb cd50996 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 | ---
language:
- en
- zh
tags:
- mlx
- object-detection
- document-layout
- reading-order
library_name: mlx-vlm
pipeline_tag: object-detection
license: apache-2.0
base_model: PaddlePaddle/PP-DocLayoutV3_safetensors
---
# PP-DocLayout V3 (MLX)
MLX conversion of [PaddlePaddle/PP-DocLayoutV3_safetensors](https://huggingface.co/PaddlePaddle/PP-DocLayoutV3_safetensors)
for document-layout detection on Apple Silicon. It predicts page regions,
class labels, and reading order; it does not transcribe text.
- **Architecture:** HGNetV2-L backbone, hybrid encoder, and deformable decoder
with a reading-order head.
- **Classes:** 25 prediction classes from the stock checkpoint.
- **Weights:** about 33M parameters, float32, 133 MB, not quantized.
## Usage
Use an [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) checkout that includes
`pp_doclayout_v3` support. From that checkout:
```sh
python -m pip install -e .
```
Replace `page.png` with a document image:
```python
from mlx_vlm.utils import get_model_path, load_model
model = load_model(get_model_path("HashNuke/pp-doclayout-v3-mlx"))
model.eval()
records = model.detect("page.png", conf=0.5)
for record in sorted(records, key=lambda item: item["reading_order"]):
print(record)
```
`get_model_path` downloads the checkpoint and returns a local path for
`load_model`. `detect` accepts an image path or a PIL image.
## Output format
Each detection contains:
- `bbox`: `[y0, x0, y1, x1]`, normalized to 0–1000 and rounded to one decimal.
- `label`: the checkpoint's class name.
- `reading_order`: a one-based rank among retained detections.
- `score`: confidence rounded to three decimals.
The returned list is in query order; sort by `reading_order` to consume it in
document order. This MLX interface returns rectangular boxes, not polygons or
segmentation masks.
The default inference recipe resizes RGB images to 1024×1024 and scales pixels
to [0, 1]. This matches IndicDocLayout's preprocessing and differs from the
stock Transformers processor's defaults. The mask-feature branch is retained
for mask-enhanced query initialization; training-only denoising weights are omitted.
## Validation
At `conf=0.5` with the same preprocessing and reading-order decoding, output
records matched fresh PyTorch reference runs on the English paper and
calendar/table samples in mlx-vlm (`examples/images/paper.png` and
`examples/images/demo_pdf1_page1.png`). Labels, reading order, and rounded
boxes/scores matched. This is sample-level output parity, not bitwise equality
of raw tensors or a detection-accuracy benchmark.
The 37-class IndicDocLayout fine-tune is available separately as the layout
stage of [IndicOCR (MLX)](https://huggingface.co/HashNuke/indic-ocr-mlx).
## License
The [upstream weights](https://huggingface.co/PaddlePaddle/PP-DocLayoutV3_safetensors)
are licensed under Apache-2.0. This conversion does not change that license.
|