--- license: apache-2.0 language: - fa library_name: paddleocr pipeline_tag: image-to-text base_model: PaddlePaddle/PP-OCRv6_medium_rec datasets: - Reza2kn/persian-handwriting-pages-3.69m - Reza2kn/persian-printed-ocr-3.5m tags: - ocr - persian - handwriting - printed-text - paddlepaddle - pp-ocrv6 --- # Bina 0.2 - RizehPizeh Bina 0.2 - RizehPizeh is a compact Persian line-recognition model fine-tuned from PP-OCRv6 medium recognition. It recognizes both handwritten and printed Persian text using a 161-symbol character dictionary and a CTC decoder. This repository preserves the earlier experimental checkpoints under `checkpoints/`. The verified release is at the repository root. ## Verified release | Held-out split | Rows | Exact lines | Exact-line accuracy | Mean normalized edit similarity | | --- | ---: | ---: | ---: | ---: | | Combined | 4,096 | 3,474 | **84.81%** | **98.81%** | | Handwriting | 2,048 | 1,758 | **85.84%** | **98.83%** | | Printed | 2,048 | 1,716 | **83.79%** | **98.79%** | Exact-line accuracy is strict: one incorrect character, digit, space, or punctuation mark makes the entire line incorrect. The independent evaluator accepted all 4,096 transforms without replacement. ## Package contents - `inference/`: exported Paddle inference graph and parameters. - `bina_text_recognition.py`: logical-order Persian inference wrapper. - `Bina-0.2-RizehPizeh.pdparams`: best training checkpoint. - `config.yml`: portable architecture and preprocessing configuration. - `persian_arabic_bina02_dict.txt`: 161-symbol character dictionary. - `MODEL_PROVENANCE.json`: revisions, counts, metrics, and SHA-256 hashes. - `eval-aligned-row-analysis-v2.json`: independent row-level evaluation. ## Inference Install the appropriate PaddlePaddle build for your hardware, then: ```bash pip install "paddleocr>=3.3.0,<4" git clone https://huggingface.co/Reza2kn/Bina-0.2-RizehPizeh cd Bina-0.2-RizehPizeh python bina_text_recognition.py path/to/line.jpg --device cpu ``` Or use it from Python: ```python from bina_text_recognition import BinaTextRecognition model = BinaTextRecognition("inference", device="cpu") for result in model.predict("line.jpg"): print(result["text"], result["score"]) ``` The exported CTC model emits Persian in visual order. Use the included wrapper to recover logical reading order while preserving Latin and numeric runs. ## Input contract The recognizer expects a horizontal, content-tight line crop: - Rotate tall detector crops 90 degrees clockwise. - Remove excessive printed-image background while retaining a small margin. - Resize to `3 x 48 x 768` with `padding: false`. - Preserve ZWNJ and Persian digits. For full-page OCR, pair this recognizer with [`PaddlePaddle/PP-OCRv6_medium_det_safetensors`](https://huggingface.co/PaddlePaddle/PP-OCRv6_medium_det_safetensors), rectify and normalize each detected crop, recognize each line, then reconstruct the page in reading order. This repository contains the recognizer, not the detector. ## Training The verified release used 169,000 balanced training lines: - 84,500 teacher-aligned handwriting crops. - 84,500 printed line images. - 4,096 page-isolated held-out evaluation lines. Training stopped automatically after validation plateaued. The best checkpoint was selected around step 47,520. Dataset revisions and preprocessing details are pinned in `MODEL_PROVENANCE.json`. ## Limitations - Persian-focused; other-language performance was not preserved or evaluated. - Full-page use requires a separate detector and reading-order reconstruction. - The handwriting split is teacher-aligned and may contain residual label noise. - Very small isolated symbols and unusually degraded crops remain difficult. ## License Apache License 2.0. See `LICENSE`.