WorldMind: Decoupled Game World Model for State-Aware NPC Behavior Paper • 2608.21439 • Published 9 days ago • 14 • 4
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures Paper • 2608.15242 • Published 12 days ago • 14 • 2
WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning Paper • 2608.22591 • Published 4 days ago • 1 • 2
ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts Paper • 2608.22510 • Published 4 days ago • 4 • 3
EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment Paper • 2608.21486 • Published 6 days ago • 1 • 2
LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks Paper • 2608.23200 • Published 3 days ago • 6 • 2
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows Paper • 2608.19741 • Published 7 days ago • 10 • 2
One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders Paper • 2606.13610 • Published 3 days ago • 5 • 2
RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling Paper • 2608.22849 • Published 3 days ago • 3 • 2
From Generation to Simulation: How Far Are World Models from Being True Simulators? Paper • 2608.23070 • Published 3 days ago • 4 • 2
Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection Paper • 2608.21455 • Published 7 days ago • 2 • 2
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection Paper • 2608.20169 • Published 3 days ago • 8 • 2
Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion Paper • 2608.19567 • Published 7 days ago • 27 • 2
GameXpert-Bench: How Far Are Coding Agents from Expert Game Development? Paper • 2608.21833 • Published 5 days ago • 15 • 1
Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports Paper • 2608.22817 • Published 3 days ago • 5 • 2
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA Paper • 2608.22856 • Published 3 days ago • 4 • 2
The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models Paper • 2608.22876 • Published 3 days ago • 29 • 2
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks Paper • 2608.23035 • Published 3 days ago • 37 • 2