Beacon: Knowing When and How to Perform Agentic Visual Reasoning Paper • 2607.28595 • Published 1 day ago • 39
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Paper • 2607.28509 • Published 1 day ago • 20
ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders Paper • 2607.21217 • Published 8 days ago • 6
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Paper • 2607.10350 • Published 16 days ago • 85
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Paper • 2607.10350 • Published 16 days ago • 85
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published 17 days ago • 102
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning Paper • 2607.07708 • Published 23 days ago • 88
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 153
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents Paper • 2606.22883 • Published Jun 22 • 37
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents Paper • 2606.22883 • Published Jun 22 • 37
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling Paper • 2606.18023 • Published Jun 16 • 210
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Paper • 2606.11042 • Published Jun 9 • 22
OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning Paper • 2606.08572 • Published Jun 7 • 14
CoVEBench: Can Video Editing Models Handle Complex Instructions? Paper • 2606.08415 • Published Jun 7 • 52
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills? Paper • 2606.01993 • Published Jun 1 • 15