| --- |
| license: apache-2.0 |
| tags: |
| - object-detection |
| - document-layout-analysis |
| - tibetan |
| - rf-detr |
| - tibla |
| pipeline_tag: object-detection |
| datasets: |
| - BDRC/TiBLAD |
| --- |
| |
| # TiBLA-RFDETR |
|
|
| **Permissive (Apache-2.0) alternative in TiBLA (Tibetan Book Layout Analysis)** — |
| a lighter, PyTorch-native RF-DETR-L detector for the page layout of modern |
| Tibetan books (headers, text area, footers, footnotes). |
|
|
| - **Base model / provenance:** [RF-DETR-L](https://github.com/roboflow/rf-detr) |
| (Roboflow, DINOv2 backbone), fine-tuned on the leak-free **v4** `tam2col` split |
| of TiBLAD. |
| - **License:** Apache-2.0. |
| - **Dataset:** [BDRC/TiBLAD](https://huggingface.co/datasets/BDRC/TiBLAD) |
| - **Paper:** [buda-base/papers](https://github.com/buda-base/papers) (`papers/2026-tibetan-book-layout`) — *arXiv link forthcoming* |
| - **Code:** [github.com/buda-base/tibla](https://github.com/buda-base/tibla) |
|
|
| ## Task |
|
|
| A **4-class** detector — `header`, `text-area`, `footer`, `footnote` — kept as |
| four classes at training time. Evaluation folds them into a **3-class canonical |
| scheme**: `header`+`footer` are combined into one `header-footer` class (matched |
| individually, merged losslessly afterwards), `text-area` is merged to a single |
| page/column envelope as a post-processing step (two boxes only on genuine |
| two-column pages), and `footnote` is left as-is. All numbers below are in that |
| canonical space, on the leak-free TiBLAD **v4** 833-page test set, unified scorer |
| (pycocotools bbox mAP 0.50:0.05:0.95; F1 by greedy IoU≥0.5 at the best-mean-F1 |
| operating point). |
|
|
| ## Inference |
|
|
| ```python |
| # pip install rfdetr |
| from rfdetr import RFDETRLarge |
| |
| model = RFDETRLarge.from_checkpoint("rfdetr_tibetan_book_layout.pth") |
| det = model.predict("page.jpg", threshold=0.26, shape=(1024, 1024)) |
| # checkpoint class ids are offset by 1 (id 0 = background): |
| # 1 header, 2 text-area, 3 footnote, 4 footer |
| ``` |
|
|
| A ready-made `infer.py` (batch, YOLO-format output, per-class thresholds) is |
| included in this repo. Recommended global operating confidence: **0.26** (the |
| best-mean-F1 point); the bundled `infer.py` also ships per-class max-F1 thresholds |
| (`header` 0.46, `text-area` 0.32, `footnote` 0.26, `footer` 0.52). |
|
|
| ## Evaluation (TiBLAD v4, 833-page test) |
|
|
| | metric | TiBLA-RTDETR | TiBLA-PP-DocLayout-L | **TiBLA-RFDETR** | |
| |---|---|---|---| |
| | license | AGPL-3.0 | Apache-2.0 | **Apache-2.0** | |
| | base model | RT-DETR-l (Ultralytics) | PP-DocLayout-L (PaddleOCR, RT-DETR-L) | **RF-DETR-L (Roboflow)** | |
| | mean F1 (canonical 3-class) | 0.959 | 0.958 | **0.927** | |
| | header-footer F1 | 0.952 | 0.951 | **0.949** | |
| | text-area F1 | 0.999 | 0.997 | **0.996** | |
| | footnote F1 | 0.925 | 0.925 | **0.835** | |
| | mean AP@0.50 | 0.974 | 0.959 | **0.925** | |
| | mean AP@[0.50:0.95] | 0.786 | 0.781 | **0.667** | |
| | shared-class mAP@[.50:.95] (DocLayNet-aligned) | 0.650 | 0.641 | **0.604** | |
| | Hidden Trespass — header/footer | 0.008 | 0.003 | **0.020** | |
| | Hidden Trespass — footnote | 0.037 | 0.037 | **0.216** | |
| | COTe (Trespass) | 0.975 (0.001) | 0.978 (0.000) | **0.974 (0.002)** | |
| | operating confidence | 0.74 | 0.68 | **0.26** | |
|
|
| *"operating confidence" is the single global best-mean-F1 confidence used for the |
| reported F1.* |
|
|
| **Hidden Trespass** = the missed peripheral (header/footer/footnote) ground-truth |
| **area** that falls inside the predicted `text-area` crop; area-based, |
| micro-averaged over the test set. Lower is better (less clutter bled into the OCR |
| region). Formal definition in the [paper](https://github.com/buda-base/papers). |
|
|
| ## Which checkpoint to pick |
|
|
| | checkpoint | license | mean F1 | shared mAP | footnote HT | |
| |---|---|---|---|---| |
| | TiBLA-RTDETR (primary) | AGPL-3.0 | 0.959 | 0.650 | 0.037 | |
| | TiBLA-PP-DocLayout-L | Apache-2.0 | 0.958 | 0.641 | 0.037 | |
| | **TiBLA-RFDETR** | Apache-2.0 | 0.927 | 0.604 | 0.216 | |
|
|
| RT-DETR-l has the top scores but its weights are AGPL-3.0 (Ultralytics). If you |
| need a permissive license, PP-DocLayout-L matches it at Apache-2.0; RF-DETR is a |
| lighter PyTorch-native Apache-2.0 option. |
|
|
| ## Citation |
|
|
| ```bibtex |
| @misc{tibla2026, |
| title = {TiBLA: Tibetan Book Layout Analysis}, |
| author = {Buddhist Digital Resource Center (BDRC)}, |
| year = {2026}, |
| howpublished = {\url{https://github.com/buda-base/tibla}}, |
| note = {arXiv link forthcoming} |
| } |
| ``` |
|
|