WazobiaVoice

image/png

Table of Contents

  1. Model Summary
  2. Model Description
  3. Inbuilt Voice Personas
  4. Speech Samples
  5. How to Use
  6. Bias, Risks, and Limitations
  7. Prohibited Uses
  8. Training
  9. Ongoing Research: VITS for Yoruba
  10. Future Improvements
  11. Citation
  12. Credits & References

Demo video generated using ememediaforge, also authored by Emmanuel Ariyo (Ememzyvisuals).

Model Summary

WazobiaVoice is a multilingual text-to-speech (TTS) and voice cloning model covering Yoruba, Hausa, Igbo, Nigerian Pidgin, and Nigerian English β€” five languages, one model, one set of weights. It is built by extending chatterbox's multilingual architecture to support four languages and orthographies it was never originally designed for, including Yoruba's tonal diacritic system, which required expanding the model's grapheme vocabulary and patching its hardcoded language-support list at both training and inference time.

WazobiaVoice ships with 13 inbuilt voice personas across all five languages β€” each individually sourced from real speakers, gender-verified against source metadata (not guessed), and quality-ranked before selection β€” plus full voice-cloning support from any reference audio clip.

Built by Axiveri, led by Emmanuel Ariyo, with a mission to bring genuinely open, commercially-usable AI voice infrastructure to African developers and startups β€” a space where the strongest existing results are locked behind non-commercial research licenses (see Ongoing Research).

Model Description

Architecture

Chatterbox's multilingual variant (ChatterboxMultilingualTTS) generates speech in two stages: a transformer backbone (T3) autoregressively predicts discrete speech tokens from input text, and a flow-matching-based generative vocoder (S3Gen) converts those tokens into a waveform. Flow-matching sits in the same generative-modeling family as diffusion β€” both learn to transform noise into structured output through an iterative process β€” and it's what gives chatterbox its voice-cloning capability: a short reference clip conditions the generation so the output waveform is shaped in that speaker's voice.

The base chatterbox architecture was pretrained on a fixed set of ~23 languages, none of which were Yoruba, Hausa, Igbo, or Nigerian Pidgin. Bringing WazobiaVoice's five target languages online required real engineering at multiple layers, not just a data swap:

  • Grapheme vocabulary expansion β€” the tokenizer's character set had to be extended to properly represent Yoruba's tonal diacritics (underdots, tone marks) as distinct, meaningful characters rather than silently normalizing or dropping them.
  • Runtime language-whitelist patch β€” the installed chatterbox package hardcodes a SUPPORTED_LANGUAGES dictionary that rejects any language_id outside its original ~23 languages at generation time, independent of what the underlying weights actually learned. WazobiaVoice's fork (see How to Use) bakes this patch in directly so the model works out of the box.
  • Multi-source, license-audited data pipeline β€” training data was assembled from a diverse collection of publicly available and properly licensed Nigerian-language speech sources, explicitly filtered to exclude anything with non-commercial restrictions, and engineered for genuine multi-speaker diversity per language rather than single-speaker bias.
  • One named exception to source-anonymity: WazobiaVoice's Hausa data is meaningfully strengthened by wazobia-tts-cc, an original Axiveri-curated dataset of high-quality Hausa political, government, and civic-affairs speech sourced from established Hausa radio stations and podcasts β€” chosen specifically because everyday political and civic discourse is a demanding, high-value register for a production voice model to get right.

How WazobiaVoice Compares

No two Nigerian-language TTS efforts have taken quite the same architectural path, and it's worth naming them directly rather than vaguely:

  • YarnGPT (SmolLM2-360M backbone + WavTokenizer) takes the same broad approach as chatterbox β€” an LLM predicting discrete audio tokens β€” and, like WazobiaVoice, produces strong Nigerian-accented English but runs into the same category of tonal/pronunciation gaps on Yoruba. Its own model card is candid about this. The hosted commercial product at yarngpt.ai produces near-perfect Yoruba output (independently verified) β€” a genuinely strong result β€” but the underlying open model still shows the gaps typical of this architecture family, suggesting the production version benefits from additional proprietary training, framework tuning, or backbone work not reflected in what's publicly released.
  • MMS-TTS (Meta) uses VITS β€” a mel-spectrogram and duration-predictor based architecture, fundamentally different from the token-autoregressive family above β€” and gets real, working Yoruba results because of it. This is strong evidence the architecture family matters here, not just data volume. The catch: Meta's MMS-TTS checkpoints are released under a non-commercial license, off-limits for production use by any startup or developer trying to build a real product.

WazobiaVoice's own Yoruba performance (see The Yoruba Tonal Problem) shows this same architectural pattern β€” which is exactly why Axiveri has an active VITS-based research track underway, aimed at closing this gap without the non-commercial restriction. See Ongoing Research.

Inbuilt Voice Personas

Every persona below was individually sourced from real speaker audio, verified for gender against the source dataset's own metadata where available, and quality-ranked before selection β€” not synthesized or guessed.

Persona Language Gender Region Traits Best Use Case
Wura Yoruba Female Lagos Social, expressive Social media content, conversational/casual content, entertainment
Bọlaji Yoruba Male Ibadan Neutral, basic General narration, everyday announcements, informational content
Elder Deji Yoruba Male β€” Elderly, native Yoruba speaker Storytelling, folklore/proverbs, cultural/heritage content, audiobooks
IfΓ© Yoruba Female β€” Teenage girl, news-style delivery Youth-oriented news reading, educational content for younger audiences
Musa Hausa Male Kano Expressive, confident Sports commentary, ads/promos, energetic announcements
Hauwa Hausa Female Kaduna Calm, professionally expressive Customer service / IVR, professional narration
Emeka Igbo Male Enugu Bold, neutrally expressive General narration, e-learning, corporate voice
Adaeze Igbo Female Imo Neutrally expressive General narration, audiobooks, virtual assistant
Tunde Nigerian Pidgin Male Lagos Calm Calm narration, wellness/meditation content
Ngozi Nigerian Pidgin Female Port Harcourt Competent, professional Customer service, business communications
John Nigerian Pidgin Male β€” Teenage boy, expressive, confident, loud Ads/promos, youth entertainment, games, energetic social content
James Nigerian English Male Lagos Island Professional Corporate narration, business communications
Amara Nigerian English Female Lagos Island Calm, friendly, expressive Friendly customer service, conversational assistant, onboarding content

Speech Samples

Native Language Samples

Persona Sample Text Audio
Wura (Yoruba, F) Ọrẹ́ mi ti dΓ© lΓ‘ti Γ¬lΓΊ αΊΉΜ€kọ́ nΓ­ Γ²nΓ­, a sΓ¬ Ε„ sọ̀rọ̀ nΓ­pa bΓ­ a αΉ£e Ε„ lọ sΓ­ ọjΓ ...
Bọlaji (Yoruba, M) Ọrẹ́ mi ti dΓ© lΓ‘ti Γ¬lΓΊ αΊΉΜ€kọ́ nΓ­ Γ²nΓ­, a sΓ¬ Ε„ sọ̀rọ̀ nΓ­pa bΓ­ a αΉ£e Ε„ lọ sΓ­ ọjΓ ...
Elder Deji (Yoruba, M) Ọrẹ́ mi ti dΓ© lΓ‘ti Γ¬lΓΊ αΊΉΜ€kọ́ nΓ­ Γ²nΓ­, a sΓ¬ Ε„ sọ̀rọ̀ nΓ­pa bΓ­ a αΉ£e Ε„ lọ sΓ­ ọjΓ ...
Ife (Yoruba, F) Ọrẹ́ mi ti dΓ© lΓ‘ti Γ¬lΓΊ αΊΉΜ€kọ́ nΓ­ Γ²nΓ­, a sΓ¬ Ε„ sọ̀rọ̀ nΓ­pa bΓ­ a αΉ£e Ε„ lọ sΓ­ ọjΓ ...
Musa (Hausa, M) Yau na tashi da safe na ci abincin safe sannan na fita zuwa kasuwa...
Hauwa (Hausa, F) Yau na tashi da safe na ci abincin safe sannan na fita zuwa kasuwa...
Emeka (Igbo, M) N'α»₯tα»₯tα»₯, m gara ahα»‹a zα»₯ta ihe oriri dα»‹ ka akwα»₯kwọ nri na azα»₯Μ€ ọhα»₯rα»₯...
Adaeze (Igbo, F) N'α»₯tα»₯tα»₯, m gara ahα»‹a zα»₯ta ihe oriri dα»‹ ka akwα»₯kwọ nri na azα»₯Μ€ ọhα»₯rα»₯...
Tunde (Pidgin, M) My guy don land from Abuja since morning, we dey gist about how we go reach market...
Ngozi (Pidgin, F) My guy don land from Abuja since morning, we dey gist about how we go reach market...
John (Pidgin, M, teen) My guy don land from Abuja since morning, we dey gist about how we go reach market...
James (English, M) My friend arrived from Lagos this morning, and we've been chatting about going to the market...
Amara (English, F) My friend arrived from Lagos this morning, and we've been chatting about going to the market...

Each persona has 2 additional sample lines in model_card_samples/ β€” this table shows one representative line per voice.

Cross-Lingual Voice Cloning

The same reference voice can speak any of the five supported languages β€” not just the one it was originally recorded in. Below, IfΓ© (a native Yoruba voice) and John (a native Pidgin voice) are shown speaking languages outside their own:

Voice (native language) Speaking Audio
IfΓ© (native: Yoruba) Hausa
IfΓ© (native: Yoruba) Igbo
IfΓ© (native: Yoruba) Nigerian Pidgin
IfΓ© (native: Yoruba) English
John (native: Pidgin) Yoruba
John (native: Pidgin) Hausa
John (native: Pidgin) Igbo
John (native: Pidgin) English

Full cross-lingual set (24 samples) available in model_card_samples/ and cross_lang_manifest.csv.

How to Use

WazobiaVoice runs on wazobiavoice-TTS, a fork of chatterbox with the language-whitelist and grapheme-vocabulary patches described in Architecture baked in directly β€” no manual patching needed.

Install

git clone https://github.com/Ememzyvisuals/wazobiavoice-TTS.git
cd wazobiavoice-TTS
bash scripts/install.sh

The install is staged deliberately (a plain pip install . will fail on this stack β€” deepfilternet and torchmetrics send pip's resolver backtracking into incompatible years-old releases). scripts/install.sh handles the correct order: system Rust toolchain for deepfilternet's extension, numpy/torch first, then the remaining audio libraries with the flags they actually need. Full reasoning is commented inline in the script itself.

Requirements: Python 3.10+, a CUDA GPU recommended (CPU works but is slow).

Quickstart

import torch
import torchaudio as ta
from wazobiavoice_tts.mtl_tts import WazobiaVoiceMultilingualTTS

device = "cuda" if torch.cuda.is_available() else "cpu"
model = WazobiaVoiceMultilingualTTS.from_pretrained(device)

wav = model.generate(
    "My guy don land from Abuja since morning, we dey gist about how we go "
    "reach market buy beans and fresh fish for evening chop.",
    language_id="pcm",
    audio_prompt_path="path/to/a_5_to_10_second_reference_clip.wav",
    exaggeration=0.55,
    cfg_weight=0.55,
)
ta.save("output.wav", wav, model.sr)

language_id accepts yo (Yoruba), ha (Hausa), ig (Igbo), pcm (Nigerian Pidgin), en (English) β€” plus the ~23 other languages inherited from the base multilingual model. audio_prompt_path is a short (5–10s) reference clip of the voice to clone; omit it to reuse whatever voice was last prepared via model.prepare_conditionals(...).

Full parameter list:

model.generate(
    text,                    # str, required
    language_id,             # str, required β€” see SUPPORTED_LANGUAGES
    audio_prompt_path=None,  # str, path to reference audio for voice cloning
    exaggeration=0.5,        # float, emotion/expressiveness intensity
    cfg_weight=0.5,          # float, classifier-free guidance weight
    temperature=0.8,
    repetition_penalty=1.2,
    min_p=0.05,
    top_p=1.0,
)

Listing supported languages:

from wazobiavoice_tts.mtl_tts import SUPPORTED_LANGUAGES
print(SUPPORTED_LANGUAGES)
# {'ar': 'Arabic', ..., 'yo': 'Yoruba', 'ha': 'Hausa', 'ig': 'Igbo', 'pcm': 'Nigerian Pidgin', 'en': 'English', ...}

Loading from a local checkpoint instead of the Hub:

model = WazobiaVoiceMultilingualTTS.from_local("path/to/checkpoint_dir", device)

Demo apps

python multilingual_app.py       # Gradio web UI, multilingual
python gradio_tts_turbo_app.py   # Gradio web UI, turbo model
python gradio_vc_app.py          # Gradio web UI, voice conversion

Full repo, install script reasoning, and license/notice files: github.com/Ememzyvisuals/wazobiavoice-TTS

Bias, Risks, and Limitations

WazobiaVoice's language coverage was built from the ground up rather than inherited, and quality is not uniform across the five supported languages:

  • Nigerian English, Igbo, and Hausa perform strongly β€” clear pronunciation, correct accent, stable delivery.
  • Nigerian Pidgin performs well, close to the English/Igbo/Hausa tier.
  • Yoruba is the clear outlier. See The Yoruba Tonal Problem below for a full technical breakdown β€” in short, the model pronounces words it saw well-represented in training reasonably well, but struggles with tonal diacritics and pronunciation on less-represented words and patterns.

The model may not capture the full range of regional accents within each language, and inbuilt persona voices reflect the specific speakers they were sourced from rather than every possible voice within a given region or demographic.

The Yoruba Tonal Problem

This deserves a direct, technical explanation rather than a vague disclaimer, because it's the single most important limitation of this model.

Yoruba is a tonal language where pitch changes the meaning of a word, marked in text through diacritics (tone marks and underdots) that are not optional decoration β€” they're load-bearing. WazobiaVoice's underlying architecture (see Architecture) generates speech through an autoregressive transformer predicting discrete tokens one at a time β€” the same broad approach used by YarnGPT. This was not a hypothesis we left untested: after the model's initial training, we ran a dedicated follow-up training pass using 36+ hours of high-diacritic-quality Yoruba audio, specifically targeting this weakness. Despite following the exact same training methodology that worked for the other four languages, quality did not improve β€” the model produced mispronunciations, unstable prosody, and in some cases outright hallucinated output. In short, the result did not come out as expected: throwing more hours of clean, well-diacritized Yoruba data at the problem did not translate into better Yoruba output, and in some cases made outputs less predictable rather than more. This is the main reason we don't treat "just fine-tune on more Yoruba data" as a real fix here β€” the data volume and quality were already there for this pass, and it still didn't move the needle.

This points to an architectural ceiling rather than a data problem: token-based autoregressive TTS appears to structurally struggle with representing tone the way Yoruba requires. This is consistent with what we observe in YarnGPT's own results, and with why architectures built specifically to solve this β€” like VITS, used by MMS-TTS β€” get meaningfully better Yoruba results by design, not luck. We're acting on this finding directly β€” see Ongoing Research.

Recommendations

Users building on WazobiaVoice for Yoruba-language use cases should test thoroughly against their own target vocabulary before production use, and should not assume uniform quality across all Yoruba text the way they reasonably could for the model's other four languages. Feedback, real-world testing reports, and contributions of additional high-quality diacritic-marked Yoruba training data are welcomed.

Prohibited Uses

In line with responsible AI practice, WazobiaVoice must not be used to:

  • Clone or impersonate a real, identifiable individual's voice without that person's explicit, informed consent
  • Generate speech for fraud, scams, impersonation, or deceptive purposes of any kind
  • Produce political disinformation, fabricated statements attributed to real public figures, or election-related deception
  • Generate hateful, harassing, or discriminatory content targeting any individual or group
  • Generate any content sexualizing or otherwise harming minors
  • Violate the license terms of any underlying dataset, architecture, or dependency this model is built on

This is not a hypothetical list β€” it reflects real decisions made during this model's development. A planned persona built from a real, identifiable public figure's voice was deliberately excluded during development specifically because informed consent wasn't in place, even though the underlying audio itself was otherwise usable.

Training

WazobiaVoice was fine-tuned via LoRA on top of chatterbox's multilingual base, across 7 epochs, on a diverse collection of publicly available and properly licensed Nigerian-language speech data β€” spanning multiple independent sources per language to ensure genuine multi-speaker diversity rather than single-speaker bias. Hausa training data is meaningfully strengthened by wazobia-tts-cc, an original Axiveri-curated dataset of high-quality political, government, and civic-affairs Hausa speech sourced from established Hausa radio stations and podcasts.

Training ran across multiple independent, interruptible cloud training sessions rather than one continuous run β€” the pipeline was built with automatic checkpoint-resume specifically so training could survive session interruptions and infrastructure limits without losing progress, picking back up exactly where it left off each time.

A dedicated Yoruba-focused continuation pass (see The Yoruba Tonal Problem) used an additional 36+ hours of high-diacritic-quality Yoruba-only audio, isolated from the other four languages to concentrate the model's full capacity on Yoruba specifically.

Ongoing Research: VITS for Yoruba

The architecture family underlying WazobiaVoice (and YarnGPT, and most public token-based Nigerian TTS efforts) has a real, demonstrated ceiling on Yoruba tonal accuracy. The architectures that have solved this β€” VITS, as used by Meta's MMS-TTS β€” are released under non-commercial licenses, unusable for any real product or startup.

Axiveri is actively researching and training a VITS-based Yoruba TTS model from scratch, using only commercially-licensed data and open training code, with the explicit goal of closing this gap for African developers and startups without the non-commercial restriction blocking every other working solution to this problem. This work is ongoing.

This research is compute-intensive and self-funded. Individual sponsors, partnerships, and donations that would help bring this to completion faster are genuinely welcomed:

Future Improvements

  • Resolve the Yoruba tonal accuracy gap via the VITS research track above
  • Expand inbuilt persona coverage (additional regional accents, additional age ranges per language)
  • Public inference API / hosted endpoint
  • Expanded voice-cloning documentation and fine-tuning guides via the GitHub fork

Citation

BibTeX:

@software{wazobiavoice2026,
  author = {Ariyo, Emmanuel},
  title  = {WazobiaVoice: Nigerian Multilingual Text-to-Speech},
  year   = {2026},
  url    = {https://github.com/Ememzyvisuals/wazobiavoice-TTS},
  note   = {Built under Axiveri. Fine-tuned from Chatterbox (Resemble AI).}
}

APA:

Ariyo, E. (2026). WazobiaVoice: Nigerian Multilingual Text-to-Speech. Built under Axiveri. https://github.com/Ememzyvisuals/wazobiavoice-TTS

Credits & References

Built and maintained by Emmanuel Ariyo (Ememzyvisuals), an independent Machine Learning and AI Engineer, under Axiveri β€” an initiative building the African AI models that African companies, startups, developers, and researchers can build on.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Axiveri/WazobiaVoice

Finetuned
(60)
this model

Space using Axiveri/WazobiaVoice 1