Instructions to use mlboydaisuke/VoxCPM2-CoreAI with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- VoxCPM
How to use mlboydaisuke/VoxCPM2-CoreAI with VoxCPM:
import soundfile as sf from voxcpm import VoxCPM model = VoxCPM.from_pretrained("mlboydaisuke/VoxCPM2-CoreAI") wav = model.generate( text="VoxCPM is an innovative end-to-end TTS model from ModelBest, designed to generate highly expressive speech.", prompt_wav_path=None, # optional: path to a prompt speech for voice cloning prompt_text=None, # optional: reference text cfg_value=2.0, # LM guidance on LocDiT, higher for better adherence to the prompt, but maybe worse inference_timesteps=10, # LocDiT inference timesteps, higher for better result, lower for fast speed normalize=True, # enable external TN tool denoise=True, # enable external Denoise tool retry_badcase=True, # enable retrying mode for some bad cases (unstoppable) retry_badcase_max_times=3, # maximum retrying times retry_badcase_ratio_threshold=6.0, # maximum length restriction for bad case detection (simple but effective), it could be adjusted for slow pace speech ) sf.write("output.wav", wav, 16000) print("saved: output.wav") - Notebooks
- Google Colab
- Kaggle
gen-cards: regenerate Use-it block
Browse files
README.md
CHANGED
|
@@ -36,7 +36,7 @@ bundles + a few host-side projections.
|
|
| 36 |
<!-- gen-cards:use-it begin id=voxcpm2-2b (managed by scripts/gen-cards β edit cards.json / QuickStart.swift, not this block) -->
|
| 37 |
## Use it
|
| 38 |
|
| 39 |
-
**New to Core AI? [Start with CoreAIKit 0.7.
|
| 40 |
|
| 41 |
β‘ **One line** β run the kit's task op on this model
|
| 42 |
(`import CoreAIOps`; no session, no model plumbing, downloads on first use):
|
|
@@ -45,13 +45,13 @@ bundles + a few host-side projections.
|
|
| 45 |
let audio = try await CoreAI.speak(text, options: .model("voxcpm2-2b"))
|
| 46 |
```
|
| 47 |
|
| 48 |
-
Every op, one shape β [Cookbook](https://github.com/john-rocky/coreai-kit/blob/0.7.
|
| 49 |
|
| 50 |
-
βΆοΈ **Run it (source)** β the [Speak runner](https://github.com/john-rocky/coreai-kit/tree/0.7.
|
| 51 |
(GUI + CLI, one app for every text-to-speech model in the catalog):
|
| 52 |
|
| 53 |
```bash
|
| 54 |
-
git clone --branch 0.7.
|
| 55 |
export DEVELOPER_DIR=/Applications/Xcode-27.0.0-RC.app/Contents/Developer
|
| 56 |
open -a /Applications/Xcode-27.0.0-RC.app coreai-kit/Examples/Speak/Speak.xcodeproj
|
| 57 |
# β Run, then pick "VoxCPM2 2B" in the model picker
|
|
@@ -73,7 +73,7 @@ let audio = try await speaker.synthesize(text)
|
|
| 73 |
// audio.samples: 48 kHz mono PCM in [-1, 1] β play it or write a WAV
|
| 74 |
```
|
| 75 |
|
| 76 |
-
The take-home is [`Examples/Speak/Sources/QuickStart.swift`](https://github.com/john-rocky/coreai-kit/blob/0.7.
|
| 77 |
β this exact code as one typed function, no UI; the CLI is an argument shell over it, and
|
| 78 |
the GUI drives the same `KitSpeaker(catalog:)` and plays the samples.
|
| 79 |
Live playback? `synthesizeStreaming(_:onChunk:)` hands you ~0.5 s chunks as they decode,
|
|
@@ -82,10 +82,10 @@ so audio starts before the whole clip exists. The WAV container is your app's te
|
|
| 82 |
|
| 83 |
**Integration checklist**
|
| 84 |
|
| 85 |
-
- SPM: `https://github.com/john-rocky/coreai-kit` (exact **0.7.
|
| 86 |
- Info.plist: none needed
|
| 87 |
- Entitlements: none needed
|
| 88 |
-
- First run downloads the model β ~4,
|
| 89 |
local cache (Application Support; progress via the `downloadProgress` callback)
|
| 90 |
- Measure in Release β Debug is ~3Γ slower on per-token host work
|
| 91 |
<!-- gen-cards:use-it end -->
|
|
|
|
| 36 |
<!-- gen-cards:use-it begin id=voxcpm2-2b (managed by scripts/gen-cards β edit cards.json / QuickStart.swift, not this block) -->
|
| 37 |
## Use it
|
| 38 |
|
| 39 |
+
**New to Core AI? [Start with CoreAIKit 0.7.3](https://github.com/john-rocky/coreai-kit#readme).** Follow its requirements and first-run steps for `qwen3-0.6b`, then open the same release's [ChatDemo](https://github.com/john-rocky/coreai-kit/tree/0.7.3/Examples/ChatDemo). The README records the tested OS/SDK and download size; model and device coverage is stated per example.
|
| 40 |
|
| 41 |
β‘ **One line** β run the kit's task op on this model
|
| 42 |
(`import CoreAIOps`; no session, no model plumbing, downloads on first use):
|
|
|
|
| 45 |
let audio = try await CoreAI.speak(text, options: .model("voxcpm2-2b"))
|
| 46 |
```
|
| 47 |
|
| 48 |
+
Every op, one shape β [Cookbook](https://github.com/john-rocky/coreai-kit/blob/0.7.3/docs/COOKBOOK.md).
|
| 49 |
|
| 50 |
+
βΆοΈ **Run it (source)** β the [Speak runner](https://github.com/john-rocky/coreai-kit/tree/0.7.3/Examples/Speak)
|
| 51 |
(GUI + CLI, one app for every text-to-speech model in the catalog):
|
| 52 |
|
| 53 |
```bash
|
| 54 |
+
git clone --branch 0.7.3 --depth 1 https://github.com/john-rocky/coreai-kit
|
| 55 |
export DEVELOPER_DIR=/Applications/Xcode-27.0.0-RC.app/Contents/Developer
|
| 56 |
open -a /Applications/Xcode-27.0.0-RC.app coreai-kit/Examples/Speak/Speak.xcodeproj
|
| 57 |
# β Run, then pick "VoxCPM2 2B" in the model picker
|
|
|
|
| 73 |
// audio.samples: 48 kHz mono PCM in [-1, 1] β play it or write a WAV
|
| 74 |
```
|
| 75 |
|
| 76 |
+
The take-home is [`Examples/Speak/Sources/QuickStart.swift`](https://github.com/john-rocky/coreai-kit/blob/0.7.3/Examples/Speak/Sources/QuickStart.swift)
|
| 77 |
β this exact code as one typed function, no UI; the CLI is an argument shell over it, and
|
| 78 |
the GUI drives the same `KitSpeaker(catalog:)` and plays the samples.
|
| 79 |
Live playback? `synthesizeStreaming(_:onChunk:)` hands you ~0.5 s chunks as they decode,
|
|
|
|
| 82 |
|
| 83 |
**Integration checklist**
|
| 84 |
|
| 85 |
+
- SPM: `https://github.com/john-rocky/coreai-kit` (exact **0.7.3**) β product **CoreAIKit**
|
| 86 |
- Info.plist: none needed
|
| 87 |
- Entitlements: none needed
|
| 88 |
+
- First run downloads the model β ~4,721 MB (Mac) / ~4,727 MB (iPhone) β then it loads from the
|
| 89 |
local cache (Application Support; progress via the `downloadProgress` callback)
|
| 90 |
- Measure in Release β Debug is ~3Γ slower on per-token host work
|
| 91 |
<!-- gen-cards:use-it end -->
|