SenseNova-U1.5-8B-MoT
Unified text-to-image and image editing model
None defined yet.
Unified text-to-image and image editing model
Action-conditioned robot manipulation video generation
Track fish in sonar video, fix it with plain English
Multilingual word and sentence alignment
Detect AI-generated images with FiSeR DINOv3 ViT-L/16
Blind image quality assessment with faithful reasoning
Assess spatial aesthetics of interior images
Few-step action-conditioned Minecraft world model
Unified text-to-image and image editing (SFT checkpoint)
Krea 2 Turbo in 4 steps via distillation LoRA
Roll a neural ocean emulator forward against GFDL-OM4
Generate human motion sequences from text prompts
4-step distilled image-to-video with SLA sparse attention
Streaming video QA from the 4 most recent frames
Chinese dialect ASR for Mandarin, Cantonese, Wu & more
Zero-shot voice cloning TTS with Audio8 0.1B
Instruction-based video editing with Qwen-Image-Edit
Turn natural-language math into Lean 4 statements
Demonstration-guided next GUI action from a screenshot
Multi-shot video with synchronized speech and foley
Generate 3D B-Rep CAD structures with HiFi-BRep
Rank document pages for a query with ConceptFormer
Danish speech-to-text with hviske-v5.3 Conformer model
Latent Shortcut based Co-Speech Gesture Generation