X-VC Voice Conversion
Zero-shot streaming voice conversion in codec space
Omni-modal image, audio and video understanding
Generate 8-panel cinematic storyboards from text prompts
Unified audio scene generation from text descriptions
Steer a 3D camera path through any still image
Generate custom music tracks from text captions and lyrics
Qwen3-VL 33B prompt conditioning as a service
Find when a described sound happens in a recording
Dense fine-grained image captioning with Qwen3-VL-4B
Proactive real-time commentary on audio-video streams