RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 5 days ago • 234
Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models Paper • 2609.39820 • Published 6 days ago • 9
Language Models that Play Chess and Explain Their Moves Paper • 2610.03695 • Published 4 days ago • 17
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 10 days ago • 322
Post-Training Leaves Behavioral Shadows on Unrelated Decisions Paper • 2609.29233 • Published 12 days ago • 272 • 2
Selecting The Most Informative Tokens in Natural Language Autoencoders Paper • 2609.37040 • Published 7 days ago • 17
Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression Paper • 2609.36322 • Published 8 days ago • 111
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 7 days ago • 114
VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models Paper • 2609.32607 • Published 10 days ago • 154