Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 1 day ago • 47
MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement Paper • 2608.14221 • Published 21 days ago • 9
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing Paper • 2607.08646 • Published Jul 9 • 3
MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis Paper • 2606.13782 • Published Jun 11 • 2
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Paper • 2602.02979 • Published Feb 3 • 1
Data Science and Technology Towards AGI Part I: Tiered Data Management Paper • 2602.09003 • Published Feb 9 • 10
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data Paper • 2505.05427 • Published May 8, 2025 • 6
AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset Paper • 2504.03612 • Published Apr 4, 2025 • 2
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Paper • 2604.13016 • Published Apr 14 • 116
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 1 day ago • 47
UltraData Collection Ultra Scale, Ultra Quality, Ultra Coverage • 15 items • Updated 3 days ago • 107