TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 7 days ago • 165
One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding Paper • 2606.30084 • Published 27 days ago • 8
stabilityai/stable-video-diffusion-img2vid-xt Image-to-Video • 2B • Updated Jul 10, 2024 • 211k • 3.36k
Macaron-A2UI: A Model for Generative UI in Personal Agents Paper • 2605.24830 • Published May 24 • 84
SOD: Step-wise On-policy Distillation for Small Language Model Agents Paper • 2605.07725 • Published May 8 • 26
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards Paper • 2605.21467 • Published May 20 • 207