Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See Paper • 2608.17744 • Published 4 days ago • 13
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents Paper • 2608.18423 • Published 4 days ago • 16
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World Paper • 2608.13546 • Published 10 days ago • 132
Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity Paper • 2608.13430 • Published 10 days ago • 13