R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets Paper • 2608.16033 • Published 25 days ago • 16
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 18 days ago • 206
PatchWorld: Gradient-Free Optimization of Executable World Models Paper • 2605.30880 • Published May 29 • 12
RubricBench: Aligning Model-Generated Rubrics with Human Standards Paper • 2603.01562 • Published Mar 2 • 64