trl-internal-testing/tiny-MuseGlimmerForConditionalGeneration Image-Text-to-Text • 6.52M • Updated 16 days ago • 69.2k • 1
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents Paper • 2608.18423 • Published 11 days ago • 20
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published 23 days ago • 109
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Paper • 2608.13417 • Published 17 days ago • 57
view article Article TutorMoments: Do AI tutors know when to help and when to hold back? allenai • 22 days ago • 31
Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements Paper • 2607.28661 • Published Jul 22 • 18
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings Paper • 2608.03994 • Published 26 days ago • 8
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 30 days ago • 97
Instella-MoE ✨ Collection Family of fully open 16B MoE LLM with 2.8B active params per token, trained on AMD Instinct™ MI300 & MI325 GPUs. • 6 items • Updated Jul 28 • 19