HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering Paper • 2607.20219 • Published Jul 22 • 1
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories Paper • 2606.02060 • Published Jun 1 • 59
Enoki Collection Enoki is a collection of models and datasets for efficient, fine-grained hallucination detection and fact extraction. • 3 items • Updated 3 days ago
Learning Selective LLM Autonomy from Copilot Feedback in Enterprise Customer Support Workflows Paper • 2604.23855 • Published Apr 26 • 2
Learning Selective LLM Autonomy from Copilot Feedback in Enterprise Customer Support Workflows Paper • 2604.23855 • Published Apr 26 • 2
The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning Paper • 2608.14229 • Published 23 days ago • 17
Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation Paper • 2606.23127 • Published Jun 22 • 26
The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar Paper • 2606.26015 • Published Jun 24 • 10