When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models Paper • 2609.19671 • Published 1 day ago • 3
Temporal Preference Optimization for Unsupervised Retrieval Paper • 2606.17664 • Published Jun 16 • 1