view reply The verification point is a really important one. We have to be careful not to accidentally optimize agents toward the shortest trajectory. An extra call that verifies state can be much more valuable than an agent stopping early.
view article Article BenchMIRT: What are LLM benchmarks actually measuring? allenai • about 1 month ago • 27
view article Article Your Agent Aced the Task. Will It Do It Again? ibm-research • 17 days ago • 118
view article Article Using Execution Traces to Evaluate AI Agent Behavior phranzia • 28 days ago • 2
view article Article Using Execution Traces to Evaluate AI Agent Behavior phranzia • 28 days ago • 2
view article Article Measuring benchmark optimization in speech recognition +5 tlebryk02, bezzam, aliceebaird, dayllon, jpc, jens-hume-ai, tzirakis • Aug 21 • 68
view article Article Inside three.ws: The Open-Source Stack That Gives AI Agents a Body, a Brain, a Wallet, and a Job three-ws • Aug 29 • 47
view article Article The Best Open Source and Open-Weight LLM Models to Run Locally in 2026 daya-shankar • May 13 • 25
view article Article We changed one line and the benchmark score moved 0.21 AUROC FINAL-Bench • Aug 22 • 15