HashNuke commited on
Commit
cd50996
·
verified ·
1 Parent(s): 9a34dcb

Clarify PP-DocLayout V3 usage

Browse files

Document native MLX loading, output coordinates, and validation scope.

Files changed (1) hide show
  1. README.md +65 -13
README.md CHANGED
@@ -6,27 +6,79 @@ tags:
6
  - mlx
7
  - object-detection
8
  - document-layout
 
9
  library_name: mlx-vlm
 
 
 
10
  ---
11
 
12
- # pp-doclayout-v3-mlx
13
 
14
- MLX port of [PaddlePaddle/PP-DocLayoutV3_safetensors](https://huggingface.co/PaddlePaddle/PP-DocLayoutV3_safetensors):
15
- RT-DETR-style document layout detector, 25 classes plus a reading-order
16
- head (33M params, float32).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
17
 
18
  ```python
19
- from mlx_vlm.utils import load_model
20
 
21
- model = load_model("HashNuke/pp-doclayout-v3-mlx")
22
  model.eval()
23
- print(model.detect("page.png", conf=0.5))
 
 
24
  ```
25
 
26
- Also usable as the layout backend for Falcon-OCR style pipelines, and
27
- the architecture twin of the 37-class IndicDocLayout fine-tune
28
- (`HashNuke/indic-ocr-mlx` `weights/layout`).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
29
 
30
- Original weights: Apache-2.0. Port code:
31
- `mlx_vlm/models/pp_doclayout_v3` in mlx-vlm. Converted with
32
- `python -m mlx_vlm.models.pp_doclayout_v3.convert --dtype float32`.
 
6
  - mlx
7
  - object-detection
8
  - document-layout
9
+ - reading-order
10
  library_name: mlx-vlm
11
+ pipeline_tag: object-detection
12
+ license: apache-2.0
13
+ base_model: PaddlePaddle/PP-DocLayoutV3_safetensors
14
  ---
15
 
16
+ # PP-DocLayout V3 (MLX)
17
 
18
+ MLX conversion of [PaddlePaddle/PP-DocLayoutV3_safetensors](https://huggingface.co/PaddlePaddle/PP-DocLayoutV3_safetensors)
19
+ for document-layout detection on Apple Silicon. It predicts page regions,
20
+ class labels, and reading order; it does not transcribe text.
21
+
22
+ - **Architecture:** HGNetV2-L backbone, hybrid encoder, and deformable decoder
23
+ with a reading-order head.
24
+ - **Classes:** 25 prediction classes from the stock checkpoint.
25
+ - **Weights:** about 33M parameters, float32, 133 MB, not quantized.
26
+
27
+ ## Usage
28
+
29
+ Use an [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) checkout that includes
30
+ `pp_doclayout_v3` support. From that checkout:
31
+
32
+ ```sh
33
+ python -m pip install -e .
34
+ ```
35
+
36
+ Replace `page.png` with a document image:
37
 
38
  ```python
39
+ from mlx_vlm.utils import get_model_path, load_model
40
 
41
+ model = load_model(get_model_path("HashNuke/pp-doclayout-v3-mlx"))
42
  model.eval()
43
+ records = model.detect("page.png", conf=0.5)
44
+ for record in sorted(records, key=lambda item: item["reading_order"]):
45
+ print(record)
46
  ```
47
 
48
+ `get_model_path` downloads the checkpoint and returns a local path for
49
+ `load_model`. `detect` accepts an image path or a PIL image.
50
+
51
+ ## Output format
52
+
53
+ Each detection contains:
54
+
55
+ - `bbox`: `[y0, x0, y1, x1]`, normalized to 0–1000 and rounded to one decimal.
56
+ - `label`: the checkpoint's class name.
57
+ - `reading_order`: a one-based rank among retained detections.
58
+ - `score`: confidence rounded to three decimals.
59
+
60
+ The returned list is in query order; sort by `reading_order` to consume it in
61
+ document order. This MLX interface returns rectangular boxes, not polygons or
62
+ segmentation masks.
63
+
64
+ The default inference recipe resizes RGB images to 1024×1024 and scales pixels
65
+ to [0, 1]. This matches IndicDocLayout's preprocessing and differs from the
66
+ stock Transformers processor's defaults. The mask-feature branch is retained
67
+ for mask-enhanced query initialization; training-only denoising weights are omitted.
68
+
69
+ ## Validation
70
+
71
+ At `conf=0.5` with the same preprocessing and reading-order decoding, output
72
+ records matched fresh PyTorch reference runs on the English paper and
73
+ calendar/table samples in mlx-vlm (`examples/images/paper.png` and
74
+ `examples/images/demo_pdf1_page1.png`). Labels, reading order, and rounded
75
+ boxes/scores matched. This is sample-level output parity, not bitwise equality
76
+ of raw tensors or a detection-accuracy benchmark.
77
+
78
+ The 37-class IndicDocLayout fine-tune is available separately as the layout
79
+ stage of [IndicOCR (MLX)](https://huggingface.co/HashNuke/indic-ocr-mlx).
80
+
81
+ ## License
82
 
83
+ The [upstream weights](https://huggingface.co/PaddlePaddle/PP-DocLayoutV3_safetensors)
84
+ are licensed under Apache-2.0. This conversion does not change that license.