WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation Paper • 2609.30221 • Published 5 days ago • 36
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 8 days ago • 156
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 19 days ago • 275
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging Paper • 2608.03316 • Published Aug 4 • 26
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published Aug 4 • 107
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper • 2607.23855 • Published Jul 26 • 28
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published Jul 13 • 60