SimWAM: A Simple World Action Model for End-to-End Autonomous Driving Paper • 2608.07468 • Published 11 days ago • 105
openai/clip-vit-large-patch14 Zero-Shot Image Classification • 0.4B • Updated Sep 15, 2023 • 6.42M • 2.07k
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 19 days ago • 303
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Paper • 2607.25659 • Published 21 days ago • 83
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 20 days ago • 139
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Paper • 2607.14614 • Published Jul 16 • 13
MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing Paper • 2607.05376 • Published Jul 6 • 10
AlayaWorld: Long-Horizon and Playable Video World Generation Paper • 2607.06291 • Published Jul 7 • 92
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Paper • 2607.04033 • Published Jul 4 • 76
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170