LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation Paper • 2609.38146 • Published 4 days ago • 10
AVA-Encoder: Towards Agent-Native Video Representation Learning Paper • 2608.12313 • Published Aug 12 • 43
What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems Paper • 2608.07565 • Published Aug 3 • 29
What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems Paper • 2608.07565 • Published Aug 3 • 29
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation Paper • 2607.24731 • Published Jul 27 • 47
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation Paper • 2607.24731 • Published Jul 27 • 47
Generalize or Detect? Towards Robust Semantic Segmentation Under Multiple Distribution Shifts Paper • 2411.03829 • Published Nov 6, 2024
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation Paper • 2508.00728 • Published Aug 1, 2025
OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps Paper • 2509.19282 • Published Sep 23, 2025 • 8