htdemucs on Core AI, float32
Browse files- .gitattributes +1 -0
- README.md +47 -0
- htdemucs_fp32.aimodel/main.hash +0 -0
- htdemucs_fp32.aimodel/main.mlirb +3 -0
- htdemucs_fp32.aimodel/metadata.json +5 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
htdemucs_fp32.aimodel/main.mlirb filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
tags:
|
| 4 |
+
- audio
|
| 5 |
+
- source-separation
|
| 6 |
+
- music
|
| 7 |
+
- demucs
|
| 8 |
+
- core-ai
|
| 9 |
+
---
|
| 10 |
+
|
| 11 |
+
# htdemucs on Core AI
|
| 12 |
+
|
| 13 |
+
Meta's [Hybrid Transformer Demucs](https://github.com/facebookresearch/demucs) (`htdemucs`: drums, bass, other, vocals) converted to a Core AI model for macOS 27 on Apple Silicon. It is the drum, bass and other separator in slurper, a stem-splitting command-line tool.
|
| 14 |
+
|
| 15 |
+
## Graph
|
| 16 |
+
|
| 17 |
+
`htdemucs_fp32.aimodel`, float32, one function `main`, fixed shapes for one 7.8 s segment at 44.1 kHz:
|
| 18 |
+
|
| 19 |
+
| | Name | Shape | Contents |
|
| 20 |
+
|---|---|---|---|
|
| 21 |
+
| in | `mix` | `[1, 2, 343980]` | stereo audio |
|
| 22 |
+
| in | `spec` | `[1, 4, 2048, 336]` | `HTDemucs._magnitude(HTDemucs._spec(mix))`: left real, left imaginary, right real, right imaginary |
|
| 23 |
+
| out | `time` | `[1, 8, 343980]` | time branch, 4 sources × 2 channels |
|
| 24 |
+
| out | `freq` | `[1, 16, 2048, 336]` | frequency branch, 4 sources × 2 channels × real/imaginary |
|
| 25 |
+
|
| 26 |
+
The graph is `HTDemucs.forward` from its normalization to just before `_mask`; both outputs are denormalized. The complex STFT stays on the host, so a caller:
|
| 27 |
+
|
| 28 |
+
1. computes `spec` as demucs does: reflect-pad by 1536 samples plus the remainder of the last hop, take a centered, normalized STFT (4096-sample periodic Hann, hop 1024), and keep bins 0..<2048 of frames 2..<338;
|
| 29 |
+
2. runs `main`;
|
| 30 |
+
3. inverts each source and channel of `freq` with `HTDemucs._ispec` (zero Nyquist bin, two zero frames each side, normalized inverse STFT, trim) and adds `time`.
|
| 31 |
+
|
| 32 |
+
For whole songs, split into 7.8 s segments overlapping by a quarter with triangular crossfades, as demucs's `apply_model` does. Source order is drums, bass, other, vocals.
|
| 33 |
+
|
| 34 |
+
float16 overflows to NaN in the frequency branch. Running two inferences at once on one loaded model corrupted the outputs, so run segments one at a time.
|
| 35 |
+
|
| 36 |
+
## Verification
|
| 37 |
+
|
| 38 |
+
- A synthetic segment through Core AI on the GPU against PyTorch `HTDemucs.forward`: 114 dB (drums), 129 dB (bass), 114 dB (other), 100 dB (vocals) SDR.
|
| 39 |
+
- A 135 s song through slurper's Swift host against PyTorch `apply_model` (no shifts, overlap 0.25): 113-120 dB SDR per stem.
|
| 40 |
+
|
| 41 |
+
## Conversion
|
| 42 |
+
|
| 43 |
+
`scripts/convert_htdemucs.py` in slurper: demucs 4.1.0, torch 2.13.0, coreai-torch 0.4.2 (coreai-core 1.0.0b2). It exports the core with `torch.export`, converts it with `TorchConverter`, and checks the Core AI output against PyTorch before saving.
|
| 44 |
+
|
| 45 |
+
## License
|
| 46 |
+
|
| 47 |
+
MIT, as are the htdemucs weights in facebookresearch/demucs.
|
htdemucs_fp32.aimodel/main.hash
ADDED
|
Binary file (32 Bytes). View file
|
|
|
htdemucs_fp32.aimodel/main.mlirb
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6fea0ebe002d1182f0ab4ad6cd933e35385e28a2cc2784e989dcee18479f5433
|
| 3 |
+
size 168673164
|
htdemucs_fp32.aimodel/metadata.json
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"producer" : "coreai-core 1.0.0b2",
|
| 3 |
+
"assetVersion" : "2.0",
|
| 4 |
+
"creationDate" : "20260914T160315Z"
|
| 5 |
+
}
|