TrevorJS commited on
Commit
5b926cd
·
verified ·
1 Parent(s): 1c3b917

htdemucs on Core AI, float32

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ htdemucs_fp32.aimodel/main.mlirb filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ tags:
4
+ - audio
5
+ - source-separation
6
+ - music
7
+ - demucs
8
+ - core-ai
9
+ ---
10
+
11
+ # htdemucs on Core AI
12
+
13
+ Meta's [Hybrid Transformer Demucs](https://github.com/facebookresearch/demucs) (`htdemucs`: drums, bass, other, vocals) converted to a Core AI model for macOS 27 on Apple Silicon. It is the drum, bass and other separator in slurper, a stem-splitting command-line tool.
14
+
15
+ ## Graph
16
+
17
+ `htdemucs_fp32.aimodel`, float32, one function `main`, fixed shapes for one 7.8 s segment at 44.1 kHz:
18
+
19
+ | | Name | Shape | Contents |
20
+ |---|---|---|---|
21
+ | in | `mix` | `[1, 2, 343980]` | stereo audio |
22
+ | in | `spec` | `[1, 4, 2048, 336]` | `HTDemucs._magnitude(HTDemucs._spec(mix))`: left real, left imaginary, right real, right imaginary |
23
+ | out | `time` | `[1, 8, 343980]` | time branch, 4 sources × 2 channels |
24
+ | out | `freq` | `[1, 16, 2048, 336]` | frequency branch, 4 sources × 2 channels × real/imaginary |
25
+
26
+ The graph is `HTDemucs.forward` from its normalization to just before `_mask`; both outputs are denormalized. The complex STFT stays on the host, so a caller:
27
+
28
+ 1. computes `spec` as demucs does: reflect-pad by 1536 samples plus the remainder of the last hop, take a centered, normalized STFT (4096-sample periodic Hann, hop 1024), and keep bins 0..<2048 of frames 2..<338;
29
+ 2. runs `main`;
30
+ 3. inverts each source and channel of `freq` with `HTDemucs._ispec` (zero Nyquist bin, two zero frames each side, normalized inverse STFT, trim) and adds `time`.
31
+
32
+ For whole songs, split into 7.8 s segments overlapping by a quarter with triangular crossfades, as demucs's `apply_model` does. Source order is drums, bass, other, vocals.
33
+
34
+ float16 overflows to NaN in the frequency branch. Running two inferences at once on one loaded model corrupted the outputs, so run segments one at a time.
35
+
36
+ ## Verification
37
+
38
+ - A synthetic segment through Core AI on the GPU against PyTorch `HTDemucs.forward`: 114 dB (drums), 129 dB (bass), 114 dB (other), 100 dB (vocals) SDR.
39
+ - A 135 s song through slurper's Swift host against PyTorch `apply_model` (no shifts, overlap 0.25): 113-120 dB SDR per stem.
40
+
41
+ ## Conversion
42
+
43
+ `scripts/convert_htdemucs.py` in slurper: demucs 4.1.0, torch 2.13.0, coreai-torch 0.4.2 (coreai-core 1.0.0b2). It exports the core with `torch.export`, converts it with `TorchConverter`, and checks the Core AI output against PyTorch before saving.
44
+
45
+ ## License
46
+
47
+ MIT, as are the htdemucs weights in facebookresearch/demucs.
htdemucs_fp32.aimodel/main.hash ADDED
Binary file (32 Bytes). View file
 
htdemucs_fp32.aimodel/main.mlirb ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6fea0ebe002d1182f0ab4ad6cd933e35385e28a2cc2784e989dcee18479f5433
3
+ size 168673164
htdemucs_fp32.aimodel/metadata.json ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ {
2
+ "producer" : "coreai-core 1.0.0b2",
3
+ "assetVersion" : "2.0",
4
+ "creationDate" : "20260914T160315Z"
5
+ }