HuggingFaceFW/fineweb-edu
Viewer
• Updated
• 3.5B • 229k
• 955
mlfoundations/dclm-baseline-1.0
Preview
• Updated
• 101k
• 254
Viewer
• Updated
• 4.48B • 67.1k
• 757
Note only multimodal data =(
Viewer
• Updated
• 48.3M • 8.95k
• 351
Viewer
• Updated
• 5.45B • 7.47k
• 473
Note Don't have directly text =(
HuggingFaceTB/issues-kaggle-notebooks
Viewer
• Updated
• 16.1M • 169
• 13
Note only 500k rows
Viewer
• Updated
• 7.89M • 14.1k
• 184
Note 1.6M rows with web-0.5-to-1.0
Locutusque/UltraTextbooks
Viewer
• Updated
• 5.52M • 1.57k
• 198
tokyotech-llm/swallow-math-v2
Viewer
• Updated
• 17.4M • 5.74k
• 26
tokyotech-llm/swallow-code-v2
Viewer
• Updated
• 147M • 174k
• 31
HuggingFaceFW/finepdfs-edu
Viewer
• Updated
• 49.5M • 4.74k
• 79
HuggingFaceTB/smollm-corpus
Viewer
• Updated
• 237M • 23.3k
• 438