card: iPhone 18 Pro measured (RTF 0.92 whole utterance / 1.53 streaming), Mac re-run, recipe verified
Browse files
README.md
CHANGED
|
@@ -23,7 +23,7 @@ Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor t
|
|
| 23 |
|
| 24 |
# Audio8-TTS-Preview-0.6b β Core AI
|
| 25 |
|
| 26 |
-
[π€ mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI](https://huggingface.co/mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI) Β· Apache-2.0 Β· base [Edge0/Audio8-TTS-Preview-0.6b](https://huggingface.co/Edge0/Audio8-TTS-Preview-0.6b)
|
| 27 |
|
| 28 |
[`Edge0/Audio8-TTS-Preview-0.6b`](https://huggingface.co/Edge0/Audio8-TTS-Preview-0.6b) (Edge0, Apache-2.0, 601M + a
|
| 29 |
337M codec) is a **DualAR text-to-speech** model of the Fish Audio S2 Pro design: a Qwen2.5-shaped **slow AR** (24
|
|
@@ -105,34 +105,43 @@ in-graph sampler, given the port's logits and the oracle's noise, makes the orac
|
|
| 105 |
| same, 15,093 fast-AR codebook draws | logits cos min 0.99805, argmax 14,705, **draw 14,711 / 15,093** |
|
| 106 |
| codec decoder on the oracle's codes, 18 utterances | wav cos min 0.99978, log-mel cos min 0.99876 |
|
| 107 |
| codec encoder (fp16w32) on the 4 reference clips, 712 frames | 691 frames exact; codebook 0 exact on every frame, every codebook β₯ 97.7 % |
|
| 108 |
-
| free run (the port's own loop, the oracle's draws) | reached eos 18 / 18; Swift host == Python engine on 1,637 / 1,688 frames (10 / 18 utterances identical end to end) |
|
| 109 |
-
| ASR round trip vs the fixture text (Fun-ASR fp32): ja CER / en WER / zh CER |
|
| 110 |
-
| speaker cosine to the reference clip (WavLM-Base-Plus-SV), 4 clone fixtures |
|
| 111 |
|
| 112 |
A miss is fp16 GPU or int8 arithmetic moving a draw that sat within a hair of the runner-up; under sampling, one flip
|
| 113 |
changes every later frame, so end-to-end identity is not the bar β the ASR and speaker rows are. They are eight to
|
| 114 |
fourteen sentences per language, our own measurement with one normalizer: the port's speech is as intelligible and as
|
| 115 |
-
speaker-faithful as the publisher's fp32 output on these fixtures, not more
|
|
|
|
| 116 |
|
| 117 |
## Speed
|
| 118 |
|
| 119 |
Measured with [`apps/Audio8Gate`](https://github.com/john-rocky/coreai-model-zoo/tree/main/apps/Audio8Gate) (the kit's `Audio8TTS` in a Release build, the 18 fixtures,
|
| 120 |
-
78 s of audio; medians; 2026-09-28
|
| 121 |
-
assets are JIT: the first load specializes them, later loads read the
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 132 |
|
| 133 |
## Use it
|
| 134 |
|
| 135 |
-
CoreAIKit `Audio8TTS`
|
| 136 |
|
| 137 |
```swift
|
| 138 |
import CoreAIKit
|
|
@@ -146,11 +155,11 @@ let cloned = try await tts.synthesize("Turn left at the second traffic light.",
|
|
| 146 |
|
| 147 |
## β¬οΈ Bundle
|
| 148 |
|
| 149 |
-
**[mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI](https://huggingface.co/mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI)** β
|
| 150 |
`audio8_dualar_int8_cl2048_w32.aimodel/` (876 MB) + `audio8_codec_decoder_fp16_t160.aimodel/` (261 MB) +
|
| 151 |
`audio8_codec_encoder_fp16w32_t216.aimodel/` (416 MB, voice registration only) + `tokenizer/`; one subtree for macOS and
|
| 152 |
iOS (JIT `.aimodel`s). Apache-2.0: LICENSE + NOTICE from the publisher's repository. Mirror:
|
| 153 |
-
`coreai-community/Audio8-TTS-Preview-0.6b-CoreAI`.
|
| 154 |
|
| 155 |
Convert yourself β [`conversion/audio8_tts/`](https://github.com/john-rocky/coreai-model-zoo/tree/main/conversion/audio8_tts): `oracle_audio8.py` β `parity_audio8.py` β
|
| 156 |
`export_audio8_frame.py --mode int8`, `export_audio8.py --part codec`, `audio8_encoder.py --frames 216 --dtype fp16w32`,
|
|
|
|
| 23 |
|
| 24 |
# Audio8-TTS-Preview-0.6b β Core AI
|
| 25 |
|
| 26 |
+
[π€ mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI](https://huggingface.co/mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI/tree/72e1c935961c786bb838e235c7839a2bd77468cc) Β· Apache-2.0 Β· base [Edge0/Audio8-TTS-Preview-0.6b](https://huggingface.co/Edge0/Audio8-TTS-Preview-0.6b)
|
| 27 |
|
| 28 |
[`Edge0/Audio8-TTS-Preview-0.6b`](https://huggingface.co/Edge0/Audio8-TTS-Preview-0.6b) (Edge0, Apache-2.0, 601M + a
|
| 29 |
337M codec) is a **DualAR text-to-speech** model of the Fish Audio S2 Pro design: a Qwen2.5-shaped **slow AR** (24
|
|
|
|
| 105 |
| same, 15,093 fast-AR codebook draws | logits cos min 0.99805, argmax 14,705, **draw 14,711 / 15,093** |
|
| 106 |
| codec decoder on the oracle's codes, 18 utterances | wav cos min 0.99978, log-mel cos min 0.99876 |
|
| 107 |
| codec encoder (fp16w32) on the 4 reference clips, 712 frames | 691 frames exact; codebook 0 exact on every frame, every codebook β₯ 97.7 % |
|
| 108 |
+
| free run (the port's own loop, the oracle's draws) | reached eos 18 / 18; Swift host (Mac) == Python engine on 1,637 / 1,688 frames (10 / 18 utterances identical end to end); the iPhone's fp16 GPU agrees with the Mac's on 385 / 1,672 frames β it diverges earlier, so its audio is gated by the rows below |
|
| 109 |
+
| ASR round trip vs the fixture text (Fun-ASR fp32): ja CER / en WER / zh CER | Mac **2.3 % / 0.9 % / 0.0 %** Β· iPhone 18 Pro **3.2 % / 0.0 % / 0.0 %** Β· the oracle's own audio 6.5 / 0.0 / 0.0 |
|
| 110 |
+
| speaker cosine to the reference clip (WavLM-Base-Plus-SV), 4 clone fixtures | Mac mean 0.855 / min 0.561 Β· iPhone 0.842 / 0.572 Β· oracle 0.854 / 0.582 |
|
| 111 |
|
| 112 |
A miss is fp16 GPU or int8 arithmetic moving a draw that sat within a hair of the runner-up; under sampling, one flip
|
| 113 |
changes every later frame, so end-to-end identity is not the bar β the ASR and speaker rows are. They are eight to
|
| 114 |
fourteen sentences per language, our own measurement with one normalizer: the port's speech is as intelligible and as
|
| 115 |
+
speaker-faithful as the publisher's fp32 output on these fixtures, not more (the Japanese "errors" are the ASR's
|
| 116 |
+
number normalization β δΈε β 30 β on both arms, plus one real word error on the phone).
|
| 117 |
|
| 118 |
## Speed
|
| 119 |
|
| 120 |
Measured with [`apps/Audio8Gate`](https://github.com/john-rocky/coreai-model-zoo/tree/main/apps/Audio8Gate) (the kit's `Audio8TTS` in a Release build, the 18 fixtures,
|
| 121 |
+
78 s of audio; medians; 2026-09-28; raw runs in [`device/`](device/), the four runs side by side in
|
| 122 |
+
[`device/tables.md`](device/tables.md)). Both assets are JIT: the first load specializes them, later loads read the
|
| 123 |
+
cache. Two paths: **streaming** (a 160-frame codec window every 32 frames, audio starts after ~1.5 s of frames) and
|
| 124 |
+
**whole utterance** (`synthesize`: the codec once at the end β one window for anything up to 7.4 s, a fifth of the codec
|
| 125 |
+
work).
|
| 126 |
+
|
| 127 |
+
| | frame (slow step + sampling + fast AR) | codec | time to first audio, streaming | RTF streaming, median / p90 | RTF whole utterance (bench, 5 s sentence) | first load (cold) β later loads | footprint |
|
| 128 |
+
|---|---|---|---|---|---|---|---|
|
| 129 |
+
| M4 Max (GPU, macOS 27 26A428, GPU lock held, another conversion running: load average 3β4) | 28 ms | 0.17 s per 160-frame window | 1.2 s | **0.81** / 0.84 | **0.70** | 5.4 s (cache 0 β 1.27 GB) β 0.65 s | 0.85β0.99 GB |
|
| 130 |
+
| iPhone 18 Pro (GPU, iOS 27 24A437, device JIT, h19p), fresh install, thermal nominal | 33β36 ms | 0.83 s per window | 2.0 s | **1.53** / 1.83 | **0.92** (nominal, after a 145 s cool-down; 0.93 on a thermally *serious* phone) | 6.1 s (cache 0 β 1.27 GB) β 0.5 s | 0.59β0.72 GB |
|
| 131 |
+
|
| 132 |
+
A 46 ms frame costs **the same ~33 ms of engine time on the phone as on the Mac** β the frame is dispatch-bound, not
|
| 133 |
+
compute-bound (a 512-slot cache, AOT compilation and Apple's composite RMSNorm/RoPE change it by 0β10 %). What
|
| 134 |
+
separates the two devices is the codec: 0.17 s per 160-frame window on the Mac, 0.83 s on the phone, so streaming
|
| 135 |
+
in 32-frame chunks β which decodes each window five times over β is real-time on the Mac and 1.5Γ real time on the
|
| 136 |
+
phone, while a whole utterance decodes on the phone at about real time. Apple's pipelined engine drives a
|
| 137 |
+
same-sized Qwen3-0.6B at 2.8 ms per token; moving the slow step onto it is the lever for a several-times faster
|
| 138 |
+
frame β a separate round. Back-to-back synthesis heats the phone: the third run in eight minutes reached *serious* and its streaming frames slowed
|
| 139 |
+
from 33 to 39 ms, but the whole-utterance bench reads the same at *serious* (0.93) as at nominal (0.92) β the frame is
|
| 140 |
+
dispatch-bound either way.
|
| 141 |
|
| 142 |
## Use it
|
| 143 |
|
| 144 |
+
CoreAIKit `Audio8TTS`; catalog id `audio8-tts-preview-0.6b` ([coreai-kit#62](https://github.com/john-rocky/coreai-kit/pull/62)):
|
| 145 |
|
| 146 |
```swift
|
| 147 |
import CoreAIKit
|
|
|
|
| 155 |
|
| 156 |
## β¬οΈ Bundle
|
| 157 |
|
| 158 |
+
**[mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI](https://huggingface.co/mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI/tree/72e1c935961c786bb838e235c7839a2bd77468cc)** (revision `72e1c935`, 2026-09-28) β
|
| 159 |
`audio8_dualar_int8_cl2048_w32.aimodel/` (876 MB) + `audio8_codec_decoder_fp16_t160.aimodel/` (261 MB) +
|
| 160 |
`audio8_codec_encoder_fp16w32_t216.aimodel/` (416 MB, voice registration only) + `tokenizer/`; one subtree for macOS and
|
| 161 |
iOS (JIT `.aimodel`s). Apache-2.0: LICENSE + NOTICE from the publisher's repository. Mirror:
|
| 162 |
+
[`coreai-community/Audio8-TTS-Preview-0.6b-CoreAI`](https://huggingface.co/coreai-community/Audio8-TTS-Preview-0.6b-CoreAI/tree/87fdb3c702b07dd9d2b25d851c4af1021e099786).
|
| 163 |
|
| 164 |
Convert yourself β [`conversion/audio8_tts/`](https://github.com/john-rocky/coreai-model-zoo/tree/main/conversion/audio8_tts): `oracle_audio8.py` β `parity_audio8.py` β
|
| 165 |
`export_audio8_frame.py --mode int8`, `export_audio8.py --part codec`, `audio8_encoder.py --frames 216 --dtype fp16w32`,
|