alexwengg's picture
Upload README.md with huggingface_hub
4557e6c verified
|
Raw
History Blame Contribute Delete
4.09 kB
---
library_name: coreml
base_model:
- ResembleAI/chatterbox
license: mit
pipeline_tag: text-to-speech
tags:
- coreml
- text-to-speech
- apple-silicon
- ios
- macos
- multilingual
language:
- en
- fr
- de
- es
- it
- pt
- pl
- tr
- ru
- nl
- cs
- ar
- zh
- ja
- hu
- ko
- hi
- da
- el
- fi
- he
- ms
- no
- sv
- sw
---
# Chatterbox Multilingual — CoreML
CoreML export of [ResembleAI/chatterbox](https://huggingface.co/ResembleAI/chatterbox)
**multilingual** (23 languages, `t3_mtl23ls_v2` + `s3gen`) for Apple platforms,
converted by [FluidInference](https://github.com/FluidInference)
(conversion toolkit: [mobius PR #89](https://github.com/FluidInference/mobius/pull/89)).
Each model ships as both `.mlpackage` (source) and compiled `.mlmodelc`.
## Models
| File | Size (fp16) | Role | Compute |
|---|---:|---|---|
| `T3-Prefill-T256-M1024-fp16` | 977 MB | Llama-520M prefill over ≤256-token context (CFG batch 2), initializes 1024-slot KV cache | CPU+GPU |
| `T3-Decode-M1024-fp16` | 977 MB | Single-step AR decode, KV cache via I/O tensors (38 ms/step) | CPU+GPU |
| `T3-Decode-M1024-fp16-stateful` | 977 MB | Single-step AR decode, KV cache in `MLState` (**16.8 ms/step**; macOS 15+/iOS 18+) | CPU+GPU |
| `Flow-N500-fp16` | 229 MB | S3Gen flow: 500-token bucket → 1000 mel frames, 10-step CFG Euler in-graph | CPU+GPU |
| `HiFT-T1000-fp16` | 40 MB | HiFTNet vocoder: mel → 24 kHz waveform (0.09 s/call) | CPU+GPU |
| `tables/tables.safetensors` | 34 MB | text/speech embedding + learned positional tables (host applies) |
| `tables/voice-default.safetensors` | 0.1 MB | precomputed built-in voice conditioning (T3 cond embeds + S3Gen ref dict) |
| `tokenizer/grapheme_mtl_merged_expanded_v1.json` | | 23-language grapheme tokenizer |
⚠️ Do **not** load the T3 packages with `.cpuOnly` — prediction hard-crashes
(also independently reported by other Chatterbox CoreML ports). Use
`.cpuAndGPU` or `.all`.
## Samples
[`samples/`](./tree/main/samples) has CoreML end-to-end renders (`e2e_*.wav`)
next to stock PyTorch renders (`baseline_*.wav`) for en/de/fr, all using the
built-in voice.
To synthesize locally without the upstream checkpoint (Apple silicon):
```bash
git clone -b feat/chatterbox-mtl-coreml https://github.com/FluidInference/mobius
cd mobius/models/tts/chatterbox/coreml
uv sync
uv run python verify/e2e_coreml.py --lang en # models auto-download from this repo
```
## Runtime boundary
The graphs cover T3 prefill/decode (with the multilingual alignment-analyzer
attention rows as outputs), the S3Gen flow, and the HiFT vocoder. The host
runtime must provide:
- text normalization + tokenization (`tokenizer/`)
- embedding prep from `tables.safetensors` (text/speech + positional; the
stock prefill context ends with **two** BOS embeds — replicate exactly)
- CFG combine `cond + w*(cond-uncond)`, repetition penalty, min-p/top-p,
sampling, EOS handling
- the `AlignmentStreamAnalyzer` heuristics, fed by the exported `align_attn`
rows (reference port: `verify/analyzer_port.py` in the conversion toolkit)
- SineGen randomness (`phase_vec`, `noise` inputs to HiFT) and CFM noise `z`
- flow bucket padding/cropping; MLState seeding from prefill KV for the
stateful decode
Voice cloning from a reference wav additionally needs the VoiceEncoder /
S3TokenizerV2 / CAMPPlus encoders, which are not converted here; voices can
be prepared offline in Python (`export-tables.py --ref-wav`) and shipped as
`voice-*.safetensors`.
## Parity (vs upstream PyTorch)
| Check | Result |
|---|---|
| T3 wrappers vs stock (fp32) | logits 3.8e-05, alignment rows exact |
| T3 CoreML fp16 | logits 2.6e-02 (range ±15), align 1.4e-03 |
| Flow CoreML fp16 | mel max 2.4e-02, mean 2.7e-03 |
| HiFT CoreML fp16 | wav max 1.7e-02, mean 2.6e-04 |
| e2e ASR round-trip (en) | exact transcript, matches PyTorch baseline |
## License
MIT, following upstream
[ResembleAI/chatterbox](https://huggingface.co/ResembleAI/chatterbox).
Upstream embeds Resemble's Perth watermarker in its Python pipeline; this
CoreML export does not include a watermarking stage.