File size: 4,349 Bytes
3efb1b0
 
 
fbce730
 
 
 
 
 
3efb1b0
fbce730
3efb1b0
 
fbce730
3efb1b0
fbce730
 
 
3efb1b0
fbce730
 
 
 
 
 
 
 
 
3efb1b0
fbce730
 
 
 
 
 
 
 
 
3efb1b0
fbce730
3efb1b0
 
fbce730
 
 
 
 
 
 
3efb1b0
 
fbce730
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3efb1b0
 
 
 
fbce730
 
 
 
 
 
3efb1b0
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
---
license: apache-2.0
tags:
  - object-detection
  - document-layout-analysis
  - tibetan
  - rf-detr
  - tibla
pipeline_tag: object-detection
datasets:
  - BDRC/TiBLAD
---

# TiBLA-RFDETR

**Permissive (Apache-2.0) alternative in TiBLA (Tibetan Book Layout Analysis)** —
a lighter, PyTorch-native RF-DETR-L detector for the page layout of modern
Tibetan books (headers, text area, footers, footnotes).

- **Base model / provenance:** [RF-DETR-L](https://github.com/roboflow/rf-detr)
  (Roboflow, DINOv2 backbone), fine-tuned on the leak-free **v4** `tam2col` split
  of TiBLAD.
- **License:** Apache-2.0.
- **Dataset:** [BDRC/TiBLAD](https://huggingface.co/datasets/BDRC/TiBLAD)
- **Paper:** [buda-base/papers](https://github.com/buda-base/papers) (`papers/2026-tibetan-book-layout`) — *arXiv link forthcoming*
- **Code:** [github.com/buda-base/tibla](https://github.com/buda-base/tibla)

## Task

A **4-class** detector — `header`, `text-area`, `footer`, `footnote` — kept as
four classes at training time. Evaluation folds them into a **3-class canonical
scheme**: `header`+`footer` are combined into one `header-footer` class (matched
individually, merged losslessly afterwards), `text-area` is merged to a single
page/column envelope as a post-processing step (two boxes only on genuine
two-column pages), and `footnote` is left as-is. All numbers below are in that
canonical space, on the leak-free TiBLAD **v4** 833-page test set, unified scorer
(pycocotools bbox mAP 0.50:0.05:0.95; F1 by greedy IoU≥0.5 at the best-mean-F1
operating point).

## Inference

```python
# pip install rfdetr
from rfdetr import RFDETRLarge

model = RFDETRLarge.from_checkpoint("rfdetr_tibetan_book_layout.pth")
det = model.predict("page.jpg", threshold=0.26, shape=(1024, 1024))
# checkpoint class ids are offset by 1 (id 0 = background):
#   1 header, 2 text-area, 3 footnote, 4 footer
```

A ready-made `infer.py` (batch, YOLO-format output, per-class thresholds) is
included in this repo. Recommended global operating confidence: **0.26** (the
best-mean-F1 point); the bundled `infer.py` also ships per-class max-F1 thresholds
(`header` 0.46, `text-area` 0.32, `footnote` 0.26, `footer` 0.52).

## Evaluation (TiBLAD v4, 833-page test)

| metric | TiBLA-RTDETR | TiBLA-PP-DocLayout-L | **TiBLA-RFDETR** |
|---|---|---|---|
| license | AGPL-3.0 | Apache-2.0 | **Apache-2.0** |
| base model | RT-DETR-l (Ultralytics) | PP-DocLayout-L (PaddleOCR, RT-DETR-L) | **RF-DETR-L (Roboflow)** |
| mean F1 (canonical 3-class) | 0.959 | 0.958 | **0.927** |
|   header-footer F1 | 0.952 | 0.951 | **0.949** |
|   text-area F1 | 0.999 | 0.997 | **0.996** |
|   footnote F1 | 0.925 | 0.925 | **0.835** |
| mean AP@0.50 | 0.974 | 0.959 | **0.925** |
| mean AP@[0.50:0.95] | 0.786 | 0.781 | **0.667** |
| shared-class mAP@[.50:.95] (DocLayNet-aligned) | 0.650 | 0.641 | **0.604** |
| Hidden Trespass — header/footer | 0.008 | 0.003 | **0.020** |
| Hidden Trespass — footnote | 0.037 | 0.037 | **0.216** |
| COTe (Trespass) | 0.975 (0.001) | 0.978 (0.000) | **0.974 (0.002)** |
| operating confidence | 0.74 | 0.68 | **0.26** |

*"operating confidence" is the single global best-mean-F1 confidence used for the
reported F1.*

**Hidden Trespass** = the missed peripheral (header/footer/footnote) ground-truth
**area** that falls inside the predicted `text-area` crop; area-based,
micro-averaged over the test set. Lower is better (less clutter bled into the OCR
region). Formal definition in the [paper](https://github.com/buda-base/papers).

## Which checkpoint to pick

| checkpoint | license | mean F1 | shared mAP | footnote HT |
|---|---|---|---|---|
| TiBLA-RTDETR (primary) | AGPL-3.0 | 0.959 | 0.650 | 0.037 |
| TiBLA-PP-DocLayout-L | Apache-2.0 | 0.958 | 0.641 | 0.037 |
| **TiBLA-RFDETR** | Apache-2.0 | 0.927 | 0.604 | 0.216 |

RT-DETR-l has the top scores but its weights are AGPL-3.0 (Ultralytics). If you
need a permissive license, PP-DocLayout-L matches it at Apache-2.0; RF-DETR is a
lighter PyTorch-native Apache-2.0 option.

## Citation

```bibtex
@misc{tibla2026,
  title        = {TiBLA: Tibetan Book Layout Analysis},
  author       = {Buddhist Digital Resource Center (BDRC)},
  year         = {2026},
  howpublished = {\url{https://github.com/buda-base/tibla}},
  note         = {arXiv link forthcoming}
}
```