--- library_name: coreml base_model: - ResembleAI/chatterbox license: mit pipeline_tag: text-to-speech tags: - coreml - text-to-speech - apple-silicon - ios - macos - multilingual language: - en - fr - de - es - it - pt - pl - tr - ru - nl - cs - ar - zh - ja - hu - ko - hi - da - el - fi - he - ms - no - sv - sw --- # Chatterbox Multilingual — CoreML CoreML export of [ResembleAI/chatterbox](https://huggingface.co/ResembleAI/chatterbox) **multilingual** (23 languages, `t3_mtl23ls_v2` + `s3gen`) for Apple platforms, converted by [FluidInference](https://github.com/FluidInference) (conversion toolkit: [mobius PR #89](https://github.com/FluidInference/mobius/pull/89)). Each model ships as both `.mlpackage` (source) and compiled `.mlmodelc`. ## Models | File | Size (fp16) | Role | Compute | |---|---:|---|---| | `T3-Prefill-T256-M1024-fp16` | 977 MB | Llama-520M prefill over ≤256-token context (CFG batch 2), initializes 1024-slot KV cache | CPU+GPU | | `T3-Decode-M1024-fp16` | 977 MB | Single-step AR decode, KV cache via I/O tensors (38 ms/step) | CPU+GPU | | `T3-Decode-M1024-fp16-stateful` | 977 MB | Single-step AR decode, KV cache in `MLState` (**16.8 ms/step**; macOS 15+/iOS 18+) | CPU+GPU | | `Flow-N500-fp16` | 229 MB | S3Gen flow: 500-token bucket → 1000 mel frames, 10-step CFG Euler in-graph | CPU+GPU | | `HiFT-T1000-fp16` | 40 MB | HiFTNet vocoder: mel → 24 kHz waveform (0.09 s/call) | CPU+GPU | | `tables/tables.safetensors` | 34 MB | text/speech embedding + learned positional tables (host applies) | | `tables/voice-default.safetensors` | 0.1 MB | precomputed built-in voice conditioning (T3 cond embeds + S3Gen ref dict) | | `tokenizer/grapheme_mtl_merged_expanded_v1.json` | | 23-language grapheme tokenizer | ⚠️ Do **not** load the T3 packages with `.cpuOnly` — prediction hard-crashes (also independently reported by other Chatterbox CoreML ports). Use `.cpuAndGPU` or `.all`. ## Samples [`samples/`](./tree/main/samples) has CoreML end-to-end renders (`e2e_*.wav`) next to stock PyTorch renders (`baseline_*.wav`) for en/de/fr, all using the built-in voice. To synthesize locally without the upstream checkpoint (Apple silicon): ```bash git clone -b feat/chatterbox-mtl-coreml https://github.com/FluidInference/mobius cd mobius/models/tts/chatterbox/coreml uv sync uv run python verify/e2e_coreml.py --lang en # models auto-download from this repo ``` ## Runtime boundary The graphs cover T3 prefill/decode (with the multilingual alignment-analyzer attention rows as outputs), the S3Gen flow, and the HiFT vocoder. The host runtime must provide: - text normalization + tokenization (`tokenizer/`) - embedding prep from `tables.safetensors` (text/speech + positional; the stock prefill context ends with **two** BOS embeds — replicate exactly) - CFG combine `cond + w*(cond-uncond)`, repetition penalty, min-p/top-p, sampling, EOS handling - the `AlignmentStreamAnalyzer` heuristics, fed by the exported `align_attn` rows (reference port: `verify/analyzer_port.py` in the conversion toolkit) - SineGen randomness (`phase_vec`, `noise` inputs to HiFT) and CFM noise `z` - flow bucket padding/cropping; MLState seeding from prefill KV for the stateful decode Voice cloning from a reference wav additionally needs the VoiceEncoder / S3TokenizerV2 / CAMPPlus encoders, which are not converted here; voices can be prepared offline in Python (`export-tables.py --ref-wav`) and shipped as `voice-*.safetensors`. ## Parity (vs upstream PyTorch) | Check | Result | |---|---| | T3 wrappers vs stock (fp32) | logits 3.8e-05, alignment rows exact | | T3 CoreML fp16 | logits 2.6e-02 (range ±15), align 1.4e-03 | | Flow CoreML fp16 | mel max 2.4e-02, mean 2.7e-03 | | HiFT CoreML fp16 | wav max 1.7e-02, mean 2.6e-04 | | e2e ASR round-trip (en) | exact transcript, matches PyTorch baseline | ## License MIT, following upstream [ResembleAI/chatterbox](https://huggingface.co/ResembleAI/chatterbox). Upstream embeds Resemble's Perth watermarker in its Python pipeline; this CoreML export does not include a watermarking stage.