PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation Paper • 2609.38597 • Published 4 days ago • 13
Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models Paper • 2603.22782 • Published Mar 24 • 19
Video2LoRA: Unified Semantic-Controlled Video Generation via Per-Reference-Video LoRA Paper • 2603.08210 • Published Apr 1
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published Aug 13 • 47
view article Article FLUX 3 Model Overview: Multimodal Flow Models for Image, Video, Audio, and Action Prediction ResterChed • Jul 24 • 25
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published Jul 18 • 140
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing Paper • 2606.26740 • Published Jun 25 • 82
UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Paper • 2606.21661 • Published Jun 19 • 29
OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data Paper • 2606.13432 • Published Jun 11 • 64