What Makes Recurrence Effective in Looped Language Models? Paper • 2609.36636 • Published 3 days ago • 10
PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond? Paper • 2609.34314 • Published 4 days ago • 4
WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation Paper • 2609.37687 • Published 3 days ago • 6
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 6 days ago • 223