Models (Text-to-Speech)
Best open-source Text-to-Speech (TTS) models โ SOTA neural voice synthesis, zero-shot cloning, multilingual & expressive speech generation.
Text-to-Speech โข Updated โข 11.5M โข โข 6.77kNote ๐ SOTA lightweight TTS (82M params). Achieves best MOS in its parameter class, outperforming models 10x larger. Ranked #1 on TTS-Arena for quality/speed ratio. ๐ Paper: https://arxiv.org/abs/2501.13067 ๐ Benchmark: TTS-Arena Leaderboard - https://huggingface.co/spaces/Pendrokar/TTS-Spaces-Arena ๐ Repo: https://github.com/hexgrad/kokoro
OpenMOSS-Team/MOSS-TTS
Text-to-Speech โข 8B โข Updated โข 6.68k โข 425Note ๐ SOTA open-source multilingual TTS by OpenMOSS. LLM-based architecture with highly expressive, high-quality speech synthesis across multiple languages. ๐ Paper: https://arxiv.org/abs/2409.03900 ๐ Benchmark: TTS-Arena Leaderboard - https://huggingface.co/spaces/Pendrokar/TTS-Spaces-Arena ๐ GitHub: https://github.com/OpenMOSS/MOSS-TTS
elbruno/Qwen3-TTS-12Hz-0.6B-Base-ONNX
Text-to-Speech โข Updated โข 9Note ๐ ONNX-optimized version of Qwen3-TTS (Alibaba). SOTA efficient multilingual TTS for edge/local inference. 12Hz codec for high-quality audio g ๐ Modelo original: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base
coqui/XTTS-v2
Text-to-Speech โข Updated โข 7.74M โข 3.75kNote ๐ SOTA en clonaciรณn de voz zero-shot multilingรผe (17 idiomas). El modelo TTS open source mรกs descargado en HF durante 2024. Mejor MOS en voice cloning entre modelos pรบblicos. ๐ Paper: https://arxiv.org/abs/2406.04904 ๐ Benchmark: VALL-E/XTTS Eval (MOS, SECS, WER) - https://paperswithcode.com/paper/xtts-mastering-multilingual-speech-synthesis ๐ GitHub: https://github.com/coqui-ai/TTS
nari-labs/Dia-1.6B
Text-to-Speech โข 2B โข Updated โข 30.5k โข โข 2.91kNote ๐ SOTA en TTS dialogal/conversacional. Primer modelo open source con sรญntesis de diรกlogos multi-hablante nativos (incluyendo risas, suspiros, emociones). Supera a ElevenLabs en MOS conversacional. ๐ Paper: https://github.com/nari-labs/dia (technical report) ๐ Benchmark: TTS-Arena Leaderboard & Seed-TTS Eval - https://huggingface.co/spaces/Pendrokar/TTS-Spaces-Arena ๐ GitHub: https://github.com/nari-labs/dia
sesame/csm-1b
Text-to-Speech โข 2B โข Updated โข 147k โข โข 2.43kNote ๐ SOTA en naturalidad conversacional (Sesame AI). Arquitectura Conversational Speech Model con contexto de larga duraciรณn. Mejor MOS en diรกlogo natural open source, evaluado como humano en blind tests. ๐ Paper: https://arxiv.org/abs/2501.13068 ๐ Benchmark: Sesame CSM Eval (MOS, UTMOS) - https://www.sesame.com/research/crossing_the_uncanny_valley_of_voice ๐ GitHub: https://github.com/SesameAILabs/csm
ResembleAI/chatterbox
Text-to-Speech โข Updated โข 1.77M โข โข 1.77kNote ๐ SOTA open-source TTS with emotion control & zero-shot voice cloning (ResembleAI). Achieves top scores on TTS-Arena with <200ms latency. Best open-source model for real-time expressive speech. ๐ Paper: https://arxiv.org/abs/2505.12212 ๐ Benchmark: TTS-Arena Leaderboard - https://huggingface.co/spaces/Pendrokar/TTS-Spaces-Arena ๐ GitHub: https://github.com/resemble-ai/chatterbox
suno/bark
Text-to-Speech โข Updated โข 37.1k โข 1.56kNote ๐ SOTA transformer-based generative TTS with non-verbal audio (laughter, sighs, music). Pioneer open-source model for highly expressive speech generation beyond standard TTS. Best open model for audio generation with prosodic richness. ๐ Paper: https://github.com/suno-ai/bark (technical blog) ๐ Benchmark: TTS-Arena + Manual MOS evaluations - https://huggingface.co/spaces/Pendrokar/TTS-Spaces-Arena ๐ GitHub: https://github.com/suno-ai/bark
SWivid/F5-TTS
Text-to-Speech โข Updated โข 789k โข 1.2kNote ๐ SOTA non-autoregressive TTS using Flow Matching with Diffusion Transformer (DiT). Trained on 100K hours multilingual data with zero-shot voice cloning, code-switching, and inference RTF of 0.15 โ fastest among diffusion-based TTS models. ๐ Paper: https://arxiv.org/abs/2410.06885 ๐ Benchmark: Seed-TTS Eval & TTS-Arena (WER, SIM-O, MOS) - https://arxiv.org/abs/2410.06885 ๐ GitHub: https://github.com/SWivid/F5-TTS
Zyphra/Zonos-v0.1-hybrid
Text-to-Speech โข 2B โข Updated โข 1.35k โข โข 1.11kNote ๐ SOTA hybrid (transformer + SSM) TTS with emotion control and zero-shot cloning. Ranks #1 on TTS-Arena for naturalness. Achieves near-human MOS (4.3+) on standard benchmarks, outperforming proprietary models. ๐ Paper: https://arxiv.org/abs/2505.02098 ๐ Benchmark: TTS-Arena & Seed-TTS Eval - https://huggingface.co/spaces/Pendrokar/TTS-Spaces-Arena ๐ GitHub: https://github.com/Zyphra/Zonos
SparkAudio/Spark-TTS-0.5B
Text-to-Speech โข Updated โข 917 โข 747Note ๐ SOTA lightweight TTS (0.5B) from SparkAudio. Top performer on Seed-TTS Eval benchmarks for zero-shot voice cloning at low parameter count. Best efficiency/quality tradeoff in open-source TTS. ๐ Paper: https://arxiv.org/abs/2503.01710 ๐ Benchmark: Seed-TTS Eval (WER, SIM-O, SIM-R) - https://arxiv.org/abs/2503.01710 ๐ GitHub: https://github.com/SparkAudio/Spark-TTS
fishaudio/fish-speech-1.5
Text-to-Speech โข Updated โข 4.05k โข 771Note ๐ SOTA multilingual TTS powered by LLM architecture (Fish Audio). Supports zero-shot voice cloning in 14+ languages with highly natural prosody. Achieves top WER and speaker similarity scores on multilingual TTS benchmarks. ๐ Paper: https://arxiv.org/abs/2411.01156 ๐ Benchmark: Seed-TTS Eval & multilingual MOS (WER, SIM-O) - https://arxiv.org/abs/2411.01156 ๐ GitHub: https://github.com/fishaudio/fish-speech
canopylabs/orpheus-3b-0.1-ft
Text-to-Speech โข 4B โข Updated โข 201k โข โข 727Note ๐ SOTA Llama-3B fine-tuned for TTS (Canopy Labs). Achieves lowest WER on Seed-TTS Eval among open-source models. Produces extremely natural speech with human-level prosody, benchmarked against GPT-4o Audio. ๐ Paper: https://canopylabs.ai/model-releases (technical blog) ๐ Benchmark: Seed-TTS Eval & UTMOS - https://huggingface.co/canopylabs/orpheus-3b-0.1-ft ๐ GitHub: https://github.com/canopylabs/orpheus-tts
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech โข 2B โข Updated โข 2.4M โข 1.92kNote ๐ SOTA multilingual TTS by Alibaba Qwen (1.7B). Custom voice variant with 12Hz codec for high-fidelity, expressive zero-shot voice cloning. Achieves near-human naturalness across multiple languages, outperforming proprietary models on standard TTS evaluations. ๐ Paper: https://arxiv.org/abs/2601.15621 ๐ Benchmark: Seed-TTS Eval & MOS (WER, SIM-O, naturalness) - https://arxiv.org/abs/2601.15621 ๐ GitHub: https://github.com/QwenLM/Qwen3-TTS
microsoft/VibeVoice-1.5B
Text-to-Speech โข 3B โข Updated โข 128k โข 2.47kNote ๐ SOTA frontier open-source TTS by Microsoft (1.5B). Designed for expressive, long-form multi-speaker podcast generation with natural turn-taking and speaker consistency. Achieves human-level naturalness on TTS-Arena. Supports English and Chinese. ๐ Paper: https://arxiv.org/abs/2508.19205 ๐ Benchmark: TTS-Arena & EmergentTTS-Eval (MOS, naturalness) - https://arxiv.org/abs/2508.19205 ๐ GitHub: https://github.com/microsoft/VibeVoice
fishaudio/s2-pro
Text-to-Speech โข 5B โข Updated โข 357k โข 1.3kNote ๐ SOTA open-source TTS by Fish Audio (S2, 5B). Features multi-speaker, multi-turn generation and instruction-following via natural language. Achieves RTF of 0.195 with time-to-first-audio below 100ms โ production-ready streaming inference. ๐ Paper: https://arxiv.org/abs/2603.08823 ๐ Benchmark: EmergentTTS-Eval & Seed-TTS Eval (WER, SIM-O, expressiveness) - https://arxiv.org/abs/2603.08823 ๐ GitHub: https://github.com/fishaudio/fish-speech
IndexTeam/IndexTTS-2
Text-to-Speech โข Updated โข 11.2k โข 780Note ๐ SOTA emotionally expressive auto-regressive zero-shot TTS (IndexTeam). Breakthrough in duration control and fine-grained emotion transfer. Achieves top scores on naturalness and speaker similarity benchmarks vs. proprietary models. ๐ Paper: https://arxiv.org/abs/2506.21619 ๐ Benchmark: Seed-TTS Eval & MOS (WER, SIM-O, emotion accuracy) - https://arxiv.org/abs/2506.21619 ๐ GitHub: https://github.com/index-tts/index-tts
bosonai/higgs-tts-2-3b-base
Text-to-Speech โข 6B โข Updated โข 540k โข 693Note ๐ SOTA expressive audio foundation model by Boson AI (6B). Pretrained on 10M+ hours of audio data with deep language and acoustic understanding. Achieves #1 win rate on EmergentTTS-Eval for expressive speech, outperforming GPT-4o Audio. ๐ Paper: https://arxiv.org/abs/2505.23009 ๐ Benchmark: EmergentTTS-Eval (win rate vs. GPT-4o Audio) - https://arxiv.org/abs/2505.23009 ๐ GitHub: https://github.com/boson-ai/higgs-audio
FunAudioLLM/Fun-CosyVoice3-0.5B-2512
Text-to-Speech โข Updated โข 48.2k โข 630Note ๐ SOTA scalable multilingual zero-shot TTS by Alibaba FunAudio (CosyVoice3, 0.5B). LLM-based with supervised semantic tokens enabling natural prosody, cross-lingual voice cloning, and fine-grained style control across multiple languages. ๐ Paper: https://arxiv.org/abs/2407.05407 ๐ Benchmark: Seed-TTS Eval & multilingual MOS (WER, SIM-O, DNSMOS) - https://arxiv.org/abs/2407.05407 ๐ GitHub: https://github.com/FunAudioLLM/CosyVoice
HKUSTAudio/Llasa-3B
Text-to-Speech โข 4B โข Updated โข 577 โข 530Note ๐ SOTA Llama-based speech synthesis scaling (HKUST, 3B). First TTS model to systematically scale both train-time and inference-time compute. Achieves SOTA on Seed-TTS Eval English & Chinese benchmarks, surpassing prior open-source models. ๐ Paper: https://arxiv.org/abs/2502.04128 ๐ Benchmark: Seed-TTS Eval (WER, SIM-O) English & Chinese - https://arxiv.org/abs/2502.04128 ๐ GitHub: https://github.com/HKUSTAudio/Llasa
myshell-ai/OpenVoiceV2
Text-to-Speech โข Updated โข 498Note ๐ SOTA zero-shot cross-lingual voice cloning by MyShell AI (OpenVoice V2). Natively supports English, Spanish, French, Chinese, Japanese & Korean with accurate tone color cloning, flexible style control (emotion, rhythm, intonation). MIT license. ๐ Paper: https://arxiv.org/abs/2312.01479 ๐ Benchmark: Speaker similarity & naturalness MOS vs. XTTS - https://arxiv.org/abs/2312.01479 ๐ GitHub: https://github.com/myshell-ai/OpenVoice
amphion/MaskGCT
Text-to-Speech โข Updated โข 351 โข 308Note ๐ SOTA zero-shot TTS with Masked Generative Codec Transformer (Amphion). Non-autoregressive approach with parallel decoding for fast, high-quality synthesis. Achieves top scores on TTS-Arena for naturalness across 6 languages (EN, ZH, FR, DE, JA, KO). ๐ Paper: https://arxiv.org/abs/2409.00750 ๐ Benchmark: TTS-Arena & SeedTTS Eval (MOS, SIM-O, WER) - https://arxiv.org/abs/2409.00750 ๐ GitHub: https://github.com/open-mmlab/Amphion