Where OCR models are actually scored. Benchmark datasets with results tables first, then per-language leaderboard Spaces. Ordered by likes.
Daniel van Strien PRO
AI & ML interests
Machine Learning Librarian
Recent Activity
upvoted an article about 4 hours ago
Building a Fast Multilingual OCR Model with Synthetic Data updated a bucket about 7 hours ago
davanstrien/trackio-bucket upvoted an article about 7 hours ago
Meet North Micro Vision: A 2.4B Native-Resolution Vision-Language ModelOrganizations
OCR: Text recognition & pipelines
Text-line and region recognisers, plus detection and recognition pipelines for documents, manga and text in photographs.
-
small-models-for-glam/kraken-ppocrv6-medium
Image-to-Text β’ Updated β’ 15 -
PaddlePaddle/PP-OCRv6_medium_rec
Image-to-Text β’ Updated β’ 64.2k β’ 29 -
kha-white/manga-ocr-base
Image-to-Text β’ Updated β’ 1.11M β’ 180 -
JustANormalTinkerer/hayai-ocr-v2
Image-to-Text β’ 0.2B β’ Updated β’ 4.78k β’ 8
OCR: Languages & scripts
OCR for particular languages and writing systems, from whole-page readers to dedicated text-line recognisers.
-
typhoon-ai/typhoon-ocr1.5-2b
Image-Text-to-Text β’ 2B β’ Updated β’ 362k β’ β’ 30 -
sbintuitions/sarashina2.2-ocr
Image-to-Text β’ 4B β’ Updated β’ 1.86k β’ 33 -
5CD-AI/Vintern-1B-v3_5
Image-Text-to-Text β’ 0.9B β’ Updated β’ 5.35k β’ 125 -
NAMAA-Space/Qari-OCR-0.4.0-VL-4B-Instruct
Image-to-Text β’ Updated β’ 908 β’ 9
Video-game gameplay datasets for agent training
Recent Hub datasets of video game gameplay (frames/video + actions, replays, trajectories) for imitation learning and game-playing agents.
Reasoning Required?
-
davanstrien/reasoning-required
Viewer β’ Updated β’ 5k β’ 111 β’ 20 -
davanstrien/ModernBERT-based-Reasoning-Required
Text Classification β’ 0.1B β’ Updated β’ 21 β’ 11 -
davanstrien/fineweb-with-reasoning-scores-and-topics
Viewer β’ Updated β’ 10k β’ 38 β’ 2 -
davanstrien/fine-reasoning-questions
Viewer β’ Updated β’ 244 β’ 69 β’ 19
Maths reasoning
Maths reasoning datasets found using https://huggingface.co/spaces/librarian-bots/huggingface-datasets-semantic-search
- Running93
Semantic Hugging Face Hub Search
π93Search Hugging Face datasets and models by meaning
-
open-r1/OpenR1-Math-220k
Viewer β’ Updated β’ 450k β’ 147k β’ 801 -
simplescaling/s1K-1.1
Viewer β’ Updated β’ 1k β’ 13.4k β’ 157 -
MU-NLPC/Calc-ape210k
Viewer β’ Updated β’ 404k β’ 2.44k β’ 27
sentence-transformers-from-synthetic-data
Example of using distilabel to generate synthetic triplets data for fine-tuning a Sentence Transformer model
-
bigcode/self-oss-instruct-sc2-exec-filter-50k
Viewer β’ Updated β’ 50.7k β’ 38.1k β’ 107 -
davanstrien/similarity-dataset-sc2-8b
Viewer β’ Updated β’ 2.32k β’ 54 β’ 6 -
davanstrien/code-prompt-similarity-model
Sentence Similarity β’ 0.1B β’ Updated β’ 80 β’ 6 -
davanstrien/abstract-wiki
Viewer β’ Updated β’ 5k β’ 18 β’ 2
haiku
πΈ This is a collection of synthetic datasets built to help improve the ability of open language models to better write haikus through the use of DPO
Probably DPO datasets
A collection of datasets that probably support DPO
-
HuggingFaceH4/ultrafeedback_binarized
Viewer β’ Updated β’ 187k β’ 24.5k β’ 348 -
mlabonne/orpo-dpo-mix-40k
Viewer β’ Updated β’ 44.2k β’ 1.56k β’ 310 -
argilla/OpenHermesPreferences
Viewer β’ Updated β’ 989k β’ 988 β’ 214 -
argilla/distilabel-capybara-dpo-7k-binarized
Viewer β’ Updated β’ 7.56k β’ 21.6k β’ 184
query-to-hub-datasets-viewer-project
OCR on the Hub
Curated OCR models for documents, languages, handwriting and text in images. Browse five collections with short practical notes.
OCR: Handwriting & archives
Handwriting, historical print and manuscript recognition. Notes on languages, text-line segmentation and transcription conventions.
-
Riksarkivet/trocr-base-handwritten-hist-swe-2
Image-to-Text β’ 0.4B β’ Updated β’ 98.3k β’ 18 -
BDRC/tibetan-ocr
Image-Text-to-Text β’ 0.8B β’ Updated β’ 9.33k β’ 4 -
isaacmg/qwen3-vl-8b-hebrew-v19a-ckpt
Image-Text-to-Text β’ Updated β’ 54 -
magistermilitum/tridis_v2_HTR_historical_manuscripts
0.6B β’ Updated β’ 380 β’ 7
OCR: Documents
Roughly ordered by recent releases, useful updates and current usage. Practical OCR and document parsing models. Reviewed September 2026.
-
ATH-MaaS/OvisOCR2
Image-Text-to-Text β’ 0.9B β’ Updated β’ 96.5k β’ β’ 463 -
StarDoc-AI/TeleOCR
Image-Text-to-Text β’ 1B β’ Updated β’ 29.7k β’ 49 -
nvidia/NVIDIA-Nemotron-Parse-2.0
Image-Text-to-Text β’ 0.9B β’ Updated β’ 108k β’ 114 -
tencent/HunyuanOCR
Image-Text-to-Text β’ 1B β’ Updated β’ 652k β’ 819
Datasets Wrapped 2025: Reasoning
The reasoning datasets that defined 2025. Part 1 of Datasets Wrapped 2025. #DatasetsWrapped2025
hub-tldr
Creating a smol model for tl;dr-ing the hub
-
davanstrien/Smol-Hub-tldr
Text Generation β’ 0.4B β’ Updated β’ 83 β’ 11 - Running93
Semantic Hugging Face Hub Search
π93Search Hugging Face datasets and models by meaning
-
davanstrien/hub-tldr-dataset-summaries-llama
Viewer β’ Updated β’ 5k β’ 28 β’ 1 -
davanstrien/hub-tldr-model-summaries-llama
Viewer β’ Updated β’ 5k β’ 44 β’ 1
synthetic-data-generation-demos
A collection of demos for various approaches to synthetic data generation
- Runtime errorAgents8
Genstruct 7B
π8 - Running on ZeroAgentsFeatured86
Instruction Synthesizer
π86Generate instruction-response pairs from text
- Running on ZeroAgentsFeatured73
Magpie
π¦73Generate and rate instruction-response pairs
- Runtime errorAgents11
Bonito
π¬11Generate task-specific instructions and responses from text
Synthetic (text) Dataset Generation
Papers about synthetic dataset generation
-
Better Synthetic Data by Retrieving and Transforming Existing Datasets
Paper β’ 2404.14361 β’ Published β’ 2 -
Generative AI for Synthetic Data Generation: Methods, Challenges and the Future
Paper β’ 2403.04190 β’ Published β’ 1 -
Best Practices and Lessons Learned on Synthetic Data for Language Models
Paper β’ 2404.07503 β’ Published β’ 32 -
A Multi-Faceted Evaluation Framework for Assessing Synthetic Data Generated by Large Language Models
Paper β’ 2404.14445 β’ Published
Historic language modeling
This collection contains models, datasets and spaces related to historic language models i.e. language models trained on historic data
-
dbmdz/bert-base-finnish-europeana-cased
Fill-Mask β’ 0.1B β’ Updated β’ 31 -
dbmdz/bert-base-historic-english-cased
Fill-Mask β’ 0.1B β’ Updated β’ 109 β’ 1 -
Livingwithmachines/erwt-year
Fill-Mask β’ Updated β’ 20 -
dbmdz/bert-base-historic-dutch-cased
Fill-Mask β’ 0.1B β’ Updated β’ 70 β’ 2
Image Preference Optimization Datasets
Datasets suitable for Image Preference Optimization based on their colum names
OCR: Leaderboards
Where OCR models are actually scored. Benchmark datasets with results tables first, then per-language leaderboard Spaces. Ordered by likes.
-
allenai/olmOCR-bench
Benchmark β’ Updated β’ 51.6k β’ 288 -
llamaindex/ParseBench
Benchmark β’ Updated β’ 169k β’ 21.9k β’ 128 -
PaddlePaddle/Real5-OmniDocBench
Benchmark β’ Updated β’ 8.82k β’ 37 - Running17
BHL OCR Leaderboard
π17OCR & VLM character error rates on 18thβ19th-c. book pages
OCR on the Hub
Curated OCR models for documents, languages, handwriting and text in images. Browse five collections with short practical notes.
OCR: Text recognition & pipelines
Text-line and region recognisers, plus detection and recognition pipelines for documents, manga and text in photographs.
-
small-models-for-glam/kraken-ppocrv6-medium
Image-to-Text β’ Updated β’ 15 -
PaddlePaddle/PP-OCRv6_medium_rec
Image-to-Text β’ Updated β’ 64.2k β’ 29 -
kha-white/manga-ocr-base
Image-to-Text β’ Updated β’ 1.11M β’ 180 -
JustANormalTinkerer/hayai-ocr-v2
Image-to-Text β’ 0.2B β’ Updated β’ 4.78k β’ 8
OCR: Handwriting & archives
Handwriting, historical print and manuscript recognition. Notes on languages, text-line segmentation and transcription conventions.
-
Riksarkivet/trocr-base-handwritten-hist-swe-2
Image-to-Text β’ 0.4B β’ Updated β’ 98.3k β’ 18 -
BDRC/tibetan-ocr
Image-Text-to-Text β’ 0.8B β’ Updated β’ 9.33k β’ 4 -
isaacmg/qwen3-vl-8b-hebrew-v19a-ckpt
Image-Text-to-Text β’ Updated β’ 54 -
magistermilitum/tridis_v2_HTR_historical_manuscripts
0.6B β’ Updated β’ 380 β’ 7
OCR: Languages & scripts
OCR for particular languages and writing systems, from whole-page readers to dedicated text-line recognisers.
-
typhoon-ai/typhoon-ocr1.5-2b
Image-Text-to-Text β’ 2B β’ Updated β’ 362k β’ β’ 30 -
sbintuitions/sarashina2.2-ocr
Image-to-Text β’ 4B β’ Updated β’ 1.86k β’ 33 -
5CD-AI/Vintern-1B-v3_5
Image-Text-to-Text β’ 0.9B β’ Updated β’ 5.35k β’ 125 -
NAMAA-Space/Qari-OCR-0.4.0-VL-4B-Instruct
Image-to-Text β’ Updated β’ 908 β’ 9
OCR: Documents
Roughly ordered by recent releases, useful updates and current usage. Practical OCR and document parsing models. Reviewed September 2026.
-
ATH-MaaS/OvisOCR2
Image-Text-to-Text β’ 0.9B β’ Updated β’ 96.5k β’ β’ 463 -
StarDoc-AI/TeleOCR
Image-Text-to-Text β’ 1B β’ Updated β’ 29.7k β’ 49 -
nvidia/NVIDIA-Nemotron-Parse-2.0
Image-Text-to-Text β’ 0.9B β’ Updated β’ 108k β’ 114 -
tencent/HunyuanOCR
Image-Text-to-Text β’ 1B β’ Updated β’ 652k β’ 819
Video-game gameplay datasets for agent training
Recent Hub datasets of video game gameplay (frames/video + actions, replays, trajectories) for imitation learning and game-playing agents.
Datasets Wrapped 2025: Reasoning
The reasoning datasets that defined 2025. Part 1 of Datasets Wrapped 2025. #DatasetsWrapped2025
Reasoning Required?
-
davanstrien/reasoning-required
Viewer β’ Updated β’ 5k β’ 111 β’ 20 -
davanstrien/ModernBERT-based-Reasoning-Required
Text Classification β’ 0.1B β’ Updated β’ 21 β’ 11 -
davanstrien/fineweb-with-reasoning-scores-and-topics
Viewer β’ Updated β’ 10k β’ 38 β’ 2 -
davanstrien/fine-reasoning-questions
Viewer β’ Updated β’ 244 β’ 69 β’ 19
hub-tldr
Creating a smol model for tl;dr-ing the hub
-
davanstrien/Smol-Hub-tldr
Text Generation β’ 0.4B β’ Updated β’ 83 β’ 11 - Running93
Semantic Hugging Face Hub Search
π93Search Hugging Face datasets and models by meaning
-
davanstrien/hub-tldr-dataset-summaries-llama
Viewer β’ Updated β’ 5k β’ 28 β’ 1 -
davanstrien/hub-tldr-model-summaries-llama
Viewer β’ Updated β’ 5k β’ 44 β’ 1
Maths reasoning
Maths reasoning datasets found using https://huggingface.co/spaces/librarian-bots/huggingface-datasets-semantic-search
- Running93
Semantic Hugging Face Hub Search
π93Search Hugging Face datasets and models by meaning
-
open-r1/OpenR1-Math-220k
Viewer β’ Updated β’ 450k β’ 147k β’ 801 -
simplescaling/s1K-1.1
Viewer β’ Updated β’ 1k β’ 13.4k β’ 157 -
MU-NLPC/Calc-ape210k
Viewer β’ Updated β’ 404k β’ 2.44k β’ 27
synthetic-data-generation-demos
A collection of demos for various approaches to synthetic data generation
- Runtime errorAgents8
Genstruct 7B
π8 - Running on ZeroAgentsFeatured86
Instruction Synthesizer
π86Generate instruction-response pairs from text
- Running on ZeroAgentsFeatured73
Magpie
π¦73Generate and rate instruction-response pairs
- Runtime errorAgents11
Bonito
π¬11Generate task-specific instructions and responses from text
sentence-transformers-from-synthetic-data
Example of using distilabel to generate synthetic triplets data for fine-tuning a Sentence Transformer model
-
bigcode/self-oss-instruct-sc2-exec-filter-50k
Viewer β’ Updated β’ 50.7k β’ 38.1k β’ 107 -
davanstrien/similarity-dataset-sc2-8b
Viewer β’ Updated β’ 2.32k β’ 54 β’ 6 -
davanstrien/code-prompt-similarity-model
Sentence Similarity β’ 0.1B β’ Updated β’ 80 β’ 6 -
davanstrien/abstract-wiki
Viewer β’ Updated β’ 5k β’ 18 β’ 2
Synthetic (text) Dataset Generation
Papers about synthetic dataset generation
-
Better Synthetic Data by Retrieving and Transforming Existing Datasets
Paper β’ 2404.14361 β’ Published β’ 2 -
Generative AI for Synthetic Data Generation: Methods, Challenges and the Future
Paper β’ 2403.04190 β’ Published β’ 1 -
Best Practices and Lessons Learned on Synthetic Data for Language Models
Paper β’ 2404.07503 β’ Published β’ 32 -
A Multi-Faceted Evaluation Framework for Assessing Synthetic Data Generated by Large Language Models
Paper β’ 2404.14445 β’ Published
haiku
πΈ This is a collection of synthetic datasets built to help improve the ability of open language models to better write haikus through the use of DPO
Historic language modeling
This collection contains models, datasets and spaces related to historic language models i.e. language models trained on historic data
-
dbmdz/bert-base-finnish-europeana-cased
Fill-Mask β’ 0.1B β’ Updated β’ 31 -
dbmdz/bert-base-historic-english-cased
Fill-Mask β’ 0.1B β’ Updated β’ 109 β’ 1 -
Livingwithmachines/erwt-year
Fill-Mask β’ Updated β’ 20 -
dbmdz/bert-base-historic-dutch-cased
Fill-Mask β’ 0.1B β’ Updated β’ 70 β’ 2
Probably DPO datasets
A collection of datasets that probably support DPO
-
HuggingFaceH4/ultrafeedback_binarized
Viewer β’ Updated β’ 187k β’ 24.5k β’ 348 -
mlabonne/orpo-dpo-mix-40k
Viewer β’ Updated β’ 44.2k β’ 1.56k β’ 310 -
argilla/OpenHermesPreferences
Viewer β’ Updated β’ 989k β’ 988 β’ 214 -
argilla/distilabel-capybara-dpo-7k-binarized
Viewer β’ Updated β’ 7.56k β’ 21.6k β’ 184
Image Preference Optimization Datasets
Datasets suitable for Image Preference Optimization based on their colum names
query-to-hub-datasets-viewer-project