Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy
Abstract
Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy compares methods along five design axes. To our knowledge, this is the first overview to systematically taxonomize joint audio-visual editing, which we map as nine edit categories spanning 28 edit types. We describe methods, datasets, and metrics for each setting and close with the open problems we view as most consequential.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Vorch-Omni: Multi-Task Orchestration of Sight and Sound (2026)
- DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution (2026)
- Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation (2026)
- Vorch-Human: Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation (2026)
- AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes (2026)
- UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos (2026)
- The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.34381 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper