DocAtlas: Long-Document Understanding as Mutable-State Interaction Paper • 2608.07527 • Published Jul 21 • 1
XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding Paper • 2608.00036 • Published Jul 21 • 1
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published Jul 16 • 145
Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners Paper • 2606.01810 • Published Jun 1
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement Paper • 2606.11926 • Published Jun 10 • 130
From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills Paper • 2605.23899 • Published May 22 • 29
SkillOpt: Executive Strategy for Self-Evolving Agent Skills Paper • 2605.23904 • Published May 22 • 264
Covering Human Action Space for Computer Use: Data Synthesis and Benchmark Paper • 2605.12501 • Published May 12 • 16
RoLD: Robot Latent Diffusion for Multi-task Policy Modeling Paper • 2403.07312 • Published Nov 4, 2024
Beyond Narrative Description: Generating Poetry from Images by Multi-Adversarial Training Paper • 1804.08473 • Published Apr 23, 2018
RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents Paper • 2602.02486 • Published Feb 2 • 20
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies Paper • 2512.05693 • Published Dec 5, 2025 • 1
InfoAgent: Advancing Autonomous Information-Seeking Agents Paper • 2509.25189 • Published Sep 29, 2025 • 15
CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment Paper • 2209.06430 • Published Sep 14, 2022
Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions Paper • 2111.10337 • Published Nov 19, 2021
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Paper • 2411.19650 • Published Nov 29, 2024
MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation Paper • 2212.09478 • Published Dec 19, 2022
Improving Diversity in Zero-Shot GAN Adaptation with Semantic Variations Paper • 2308.10554 • Published Aug 21, 2023
SINC: Self-Supervised In-Context Learning for Vision-Language Tasks Paper • 2307.07742 • Published Jul 15, 2023