Convert source voice to match a reference voice
Convert speech to sound like another voice
Transcribe WAV audio into word‑level timestamps
Generate word‑level timestamps from uploaded WAV audio
Decode Whisper encoder output into timed subtitles
Extract vocals from any song in seconds