Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published 20 days ago • 138
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters Paper • 2602.10604 • Published Feb 11 • 201
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters Paper • 2602.10604 • Published Feb 11 • 201
Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction Paper • 2511.23476 • Published Nov 28, 2025
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning Paper • 2601.05593 • Published Jan 9 • 87
PaCoRe Collection Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning • 4 items • Updated Jan 16 • 12