mlboydaisuke commited on
Commit
90fecbe
Β·
verified Β·
1 Parent(s): 72e1c93

card: iPhone 18 Pro measured (RTF 0.92 whole utterance / 1.53 streaming), Mac re-run, recipe verified

Browse files
Files changed (1) hide show
  1. README.md +29 -20
README.md CHANGED
@@ -23,7 +23,7 @@ Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor t
23
 
24
  # Audio8-TTS-Preview-0.6b β€” Core AI
25
 
26
- [πŸ€— mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI](https://huggingface.co/mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI) Β· Apache-2.0 Β· base [Edge0/Audio8-TTS-Preview-0.6b](https://huggingface.co/Edge0/Audio8-TTS-Preview-0.6b)
27
 
28
  [`Edge0/Audio8-TTS-Preview-0.6b`](https://huggingface.co/Edge0/Audio8-TTS-Preview-0.6b) (Edge0, Apache-2.0, 601M + a
29
  337M codec) is a **DualAR text-to-speech** model of the Fish Audio S2 Pro design: a Qwen2.5-shaped **slow AR** (24
@@ -105,34 +105,43 @@ in-graph sampler, given the port's logits and the oracle's noise, makes the orac
105
  | same, 15,093 fast-AR codebook draws | logits cos min 0.99805, argmax 14,705, **draw 14,711 / 15,093** |
106
  | codec decoder on the oracle's codes, 18 utterances | wav cos min 0.99978, log-mel cos min 0.99876 |
107
  | codec encoder (fp16w32) on the 4 reference clips, 712 frames | 691 frames exact; codebook 0 exact on every frame, every codebook β‰₯ 97.7 % |
108
- | free run (the port's own loop, the oracle's draws) | reached eos 18 / 18; Swift host == Python engine on 1,637 / 1,688 frames (10 / 18 utterances identical end to end) |
109
- | ASR round trip vs the fixture text (Fun-ASR fp32): ja CER / en WER / zh CER | port **2.3 % / 0.9 % / 0.0 %** Β· the oracle's own audio 6.5 / 0.0 / 0.0 |
110
- | speaker cosine to the reference clip (WavLM-Base-Plus-SV), 4 clone fixtures | port mean 0.855 / min 0.561 Β· oracle 0.854 / 0.582 |
111
 
112
  A miss is fp16 GPU or int8 arithmetic moving a draw that sat within a hair of the runner-up; under sampling, one flip
113
  changes every later frame, so end-to-end identity is not the bar β€” the ASR and speaker rows are. They are eight to
114
  fourteen sentences per language, our own measurement with one normalizer: the port's speech is as intelligible and as
115
- speaker-faithful as the publisher's fp32 output on these fixtures, not more.
 
116
 
117
  ## Speed
118
 
119
  Measured with [`apps/Audio8Gate`](https://github.com/john-rocky/coreai-model-zoo/tree/main/apps/Audio8Gate) (the kit's `Audio8TTS` in a Release build, the 18 fixtures,
120
- 78 s of audio; medians; 2026-09-28) **while another model conversion ran on the same Mac (load average 5–7)**. Both
121
- assets are JIT: the first load specializes them, later loads read the cache.
122
-
123
- | | frame (slow step + sampling + fast AR) | codec (160-frame window) | time to first audio (32-frame chunk) | RTF median / p90 | first load (cold) β†’ later loads | footprint |
124
- |---|---|---|---|---|---|---|
125
- | M4 Max (GPU, macOS 27 26A428, GPU lock held) | 36–40 ms | ~0.7 s per utterance | 1.5 s | **1.05** / 1.09 (bench on one 5 s sentence: 0.99) | 5.4 s (cache 0 β†’ 1.27 GB) β†’ 0.8 s | 0.85–0.95 GB |
126
- | iPhone 18 Pro | not yet measured (the device was held by another lane on 2026-09-28) | | | | | |
127
-
128
- A 46 ms frame costs ~36 ms of engine time, so the Mac synthesizes at about real time and streams the first 1.5 s of
129
- audio after 1.5 s. The per-frame cost is the raw runtime path's per-layer dispatch (a 512-slot cache, AOT compilation
130
- and Apple's composite RMSNorm/RoPE change it by 0–10 %); Apple's pipelined engine drives a same-sized Qwen3-0.6B at
131
- 2.8 ms per token, and moving the slow step onto it is the lever for a several-times faster frame β€” a separate round.
 
 
 
 
 
 
 
 
132
 
133
  ## Use it
134
 
135
- CoreAIKit `Audio8TTS` (catalog enrollment pending):
136
 
137
  ```swift
138
  import CoreAIKit
@@ -146,11 +155,11 @@ let cloned = try await tts.synthesize("Turn left at the second traffic light.",
146
 
147
  ## ⬇️ Bundle
148
 
149
- **[mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI](https://huggingface.co/mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI)** β€”
150
  `audio8_dualar_int8_cl2048_w32.aimodel/` (876 MB) + `audio8_codec_decoder_fp16_t160.aimodel/` (261 MB) +
151
  `audio8_codec_encoder_fp16w32_t216.aimodel/` (416 MB, voice registration only) + `tokenizer/`; one subtree for macOS and
152
  iOS (JIT `.aimodel`s). Apache-2.0: LICENSE + NOTICE from the publisher's repository. Mirror:
153
- `coreai-community/Audio8-TTS-Preview-0.6b-CoreAI`.
154
 
155
  Convert yourself β€” [`conversion/audio8_tts/`](https://github.com/john-rocky/coreai-model-zoo/tree/main/conversion/audio8_tts): `oracle_audio8.py` β†’ `parity_audio8.py` β†’
156
  `export_audio8_frame.py --mode int8`, `export_audio8.py --part codec`, `audio8_encoder.py --frames 216 --dtype fp16w32`,
 
23
 
24
  # Audio8-TTS-Preview-0.6b β€” Core AI
25
 
26
+ [πŸ€— mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI](https://huggingface.co/mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI/tree/72e1c935961c786bb838e235c7839a2bd77468cc) Β· Apache-2.0 Β· base [Edge0/Audio8-TTS-Preview-0.6b](https://huggingface.co/Edge0/Audio8-TTS-Preview-0.6b)
27
 
28
  [`Edge0/Audio8-TTS-Preview-0.6b`](https://huggingface.co/Edge0/Audio8-TTS-Preview-0.6b) (Edge0, Apache-2.0, 601M + a
29
  337M codec) is a **DualAR text-to-speech** model of the Fish Audio S2 Pro design: a Qwen2.5-shaped **slow AR** (24
 
105
  | same, 15,093 fast-AR codebook draws | logits cos min 0.99805, argmax 14,705, **draw 14,711 / 15,093** |
106
  | codec decoder on the oracle's codes, 18 utterances | wav cos min 0.99978, log-mel cos min 0.99876 |
107
  | codec encoder (fp16w32) on the 4 reference clips, 712 frames | 691 frames exact; codebook 0 exact on every frame, every codebook β‰₯ 97.7 % |
108
+ | free run (the port's own loop, the oracle's draws) | reached eos 18 / 18; Swift host (Mac) == Python engine on 1,637 / 1,688 frames (10 / 18 utterances identical end to end); the iPhone's fp16 GPU agrees with the Mac's on 385 / 1,672 frames β€” it diverges earlier, so its audio is gated by the rows below |
109
+ | ASR round trip vs the fixture text (Fun-ASR fp32): ja CER / en WER / zh CER | Mac **2.3 % / 0.9 % / 0.0 %** Β· iPhone 18 Pro **3.2 % / 0.0 % / 0.0 %** Β· the oracle's own audio 6.5 / 0.0 / 0.0 |
110
+ | speaker cosine to the reference clip (WavLM-Base-Plus-SV), 4 clone fixtures | Mac mean 0.855 / min 0.561 Β· iPhone 0.842 / 0.572 Β· oracle 0.854 / 0.582 |
111
 
112
  A miss is fp16 GPU or int8 arithmetic moving a draw that sat within a hair of the runner-up; under sampling, one flip
113
  changes every later frame, so end-to-end identity is not the bar β€” the ASR and speaker rows are. They are eight to
114
  fourteen sentences per language, our own measurement with one normalizer: the port's speech is as intelligible and as
115
+ speaker-faithful as the publisher's fp32 output on these fixtures, not more (the Japanese "errors" are the ASR's
116
+ number normalization β€” 三十 β†’ 30 β€” on both arms, plus one real word error on the phone).
117
 
118
  ## Speed
119
 
120
  Measured with [`apps/Audio8Gate`](https://github.com/john-rocky/coreai-model-zoo/tree/main/apps/Audio8Gate) (the kit's `Audio8TTS` in a Release build, the 18 fixtures,
121
+ 78 s of audio; medians; 2026-09-28; raw runs in [`device/`](device/), the four runs side by side in
122
+ [`device/tables.md`](device/tables.md)). Both assets are JIT: the first load specializes them, later loads read the
123
+ cache. Two paths: **streaming** (a 160-frame codec window every 32 frames, audio starts after ~1.5 s of frames) and
124
+ **whole utterance** (`synthesize`: the codec once at the end β€” one window for anything up to 7.4 s, a fifth of the codec
125
+ work).
126
+
127
+ | | frame (slow step + sampling + fast AR) | codec | time to first audio, streaming | RTF streaming, median / p90 | RTF whole utterance (bench, 5 s sentence) | first load (cold) β†’ later loads | footprint |
128
+ |---|---|---|---|---|---|---|---|
129
+ | M4 Max (GPU, macOS 27 26A428, GPU lock held, another conversion running: load average 3–4) | 28 ms | 0.17 s per 160-frame window | 1.2 s | **0.81** / 0.84 | **0.70** | 5.4 s (cache 0 β†’ 1.27 GB) β†’ 0.65 s | 0.85–0.99 GB |
130
+ | iPhone 18 Pro (GPU, iOS 27 24A437, device JIT, h19p), fresh install, thermal nominal | 33–36 ms | 0.83 s per window | 2.0 s | **1.53** / 1.83 | **0.92** (nominal, after a 145 s cool-down; 0.93 on a thermally *serious* phone) | 6.1 s (cache 0 β†’ 1.27 GB) β†’ 0.5 s | 0.59–0.72 GB |
131
+
132
+ A 46 ms frame costs **the same ~33 ms of engine time on the phone as on the Mac** β€” the frame is dispatch-bound, not
133
+ compute-bound (a 512-slot cache, AOT compilation and Apple's composite RMSNorm/RoPE change it by 0–10 %). What
134
+ separates the two devices is the codec: 0.17 s per 160-frame window on the Mac, 0.83 s on the phone, so streaming
135
+ in 32-frame chunks β€” which decodes each window five times over β€” is real-time on the Mac and 1.5Γ— real time on the
136
+ phone, while a whole utterance decodes on the phone at about real time. Apple's pipelined engine drives a
137
+ same-sized Qwen3-0.6B at 2.8 ms per token; moving the slow step onto it is the lever for a several-times faster
138
+ frame β€” a separate round. Back-to-back synthesis heats the phone: the third run in eight minutes reached *serious* and its streaming frames slowed
139
+ from 33 to 39 ms, but the whole-utterance bench reads the same at *serious* (0.93) as at nominal (0.92) β€” the frame is
140
+ dispatch-bound either way.
141
 
142
  ## Use it
143
 
144
+ CoreAIKit `Audio8TTS`; catalog id `audio8-tts-preview-0.6b` ([coreai-kit#62](https://github.com/john-rocky/coreai-kit/pull/62)):
145
 
146
  ```swift
147
  import CoreAIKit
 
155
 
156
  ## ⬇️ Bundle
157
 
158
+ **[mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI](https://huggingface.co/mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI/tree/72e1c935961c786bb838e235c7839a2bd77468cc)** (revision `72e1c935`, 2026-09-28) β€”
159
  `audio8_dualar_int8_cl2048_w32.aimodel/` (876 MB) + `audio8_codec_decoder_fp16_t160.aimodel/` (261 MB) +
160
  `audio8_codec_encoder_fp16w32_t216.aimodel/` (416 MB, voice registration only) + `tokenizer/`; one subtree for macOS and
161
  iOS (JIT `.aimodel`s). Apache-2.0: LICENSE + NOTICE from the publisher's repository. Mirror:
162
+ [`coreai-community/Audio8-TTS-Preview-0.6b-CoreAI`](https://huggingface.co/coreai-community/Audio8-TTS-Preview-0.6b-CoreAI/tree/87fdb3c702b07dd9d2b25d851c4af1021e099786).
163
 
164
  Convert yourself β€” [`conversion/audio8_tts/`](https://github.com/john-rocky/coreai-model-zoo/tree/main/conversion/audio8_tts): `oracle_audio8.py` β†’ `parity_audio8.py` β†’
165
  `export_audio8_frame.py --mode int8`, `export_audio8.py --part codec`, `audio8_encoder.py --frames 216 --dtype fp16w32`,