coder543 commited on
Commit
a2bf033
·
verified ·
1 Parent(s): 3f3de01

Publish qualified 45-second and four-minute quality bundle

Browse files
README.md CHANGED
@@ -14,73 +14,85 @@ See LICENSE and NOTICE. This conversion is not endorsed by NVIDIA.
14
  Choose one self-contained bundle:
15
 
16
  - `fast/`: full finite Fourier relative attention, 192 encoder positions,
17
- fixed 15-second chunks, up to 128 independent chunks in the FP16 ANE decoder.
18
- - `quality/`: full sinusoidal relative attention, 3,776 encoder positions,
19
- exact tiled convolution subsampling, fixed 300-second chunks and an FP32 CPU
20
- decoder batching up to four independent chunks.
21
-
22
- Both use W8A16 encoder weights, retaining selected sensitive projections in
 
 
23
  FP16, and four concurrent encoder requests. Neither truncates attention within
24
- its chunk or drops input audio. `runtime.json` records the qualified scheduling
25
- policy. The two bundles trade speed for longer context; the labels describe the
26
- measured JFK result rather than a universal quality ranking.
 
27
 
28
  Requires physical Apple silicon on macOS 27 or iOS 27 and a model-specific host
29
  runtime. Graphs accept model features and recurrent states, not audio files.
30
  The host must implement the frontend, greedy TDT loop and tokenizer described
31
- by metadata/sidecars. Weights and recurrent states are shared across static
32
- functions; recurrent histories remain independent between utterances.
33
- Vocabulary: 1024 nonblank tokens, blank ID
34
- 1024, durations 0/1/2/3/4 and at most ten symbols per frame.
35
- For the larger V3 vocabulary the batched graph returns `(partition, local)` in
36
- two FP16 channels; reconstruct a token as `partition * 2048 + local`.
37
- Global IDs above 2,048 must not be transported as a single FP16 value.
38
- The scalar FP32 decoder returns a global token ID.
39
 
40
  ## Measured performance
41
 
42
  M3 MacBook Air (16 GB), macOS 27 build 26A428. Medians after warmup;
43
- frontend, encoder and decoding are included, while preparation, file I/O and
44
- chunk planning are excluded. The audio is JFK's **“We choose to go to the Moon”**
45
- speech: a **20-second excerpt** and the **18-minute 15-second recording**.
46
- The quality short-clip test intentionally uses the unmodified five-minute shape.
47
 
48
  | Bundle | Audio duration | Transcription | Audio / elapsed |
49
  | --- | ---: | ---: | ---: |
50
  | fast | 20 seconds | 0.150 s | 133.4× |
51
  | fast | 18 min 15 s | 1.876 s | 584.0× |
52
- | quality | 20 seconds | 1.336 s | 15.0× |
53
- | quality | 18 min 15 s | 7.044 s | 155.5× |
 
 
 
 
54
 
55
  Whisper-normalized WER against the supplied long-recording reference:
56
 
57
- | Bundle | Word errors | WER |
58
  | --- | ---: | ---: |
59
  | fast | 79/2220 | 3.56% |
60
- | quality | 44/2220 | 1.98% |
61
-
62
- This is one English recording, not a general or multilingual accuracy evaluation.
63
- Native quantization and FP16 decoder operation order can change token decisions.
64
- The reported full-recording runs retain all audio, with 74 fast chunks or four
65
- quality chunks. These are distinct contexts and execution policies, not a
66
- controlled kernel-only comparison.
 
 
 
 
 
 
 
 
 
 
 
67
 
68
  ## Device specialization and loading
69
 
70
- Source `.aimodel` files specialize on the device. On the same Mac, preparation
71
- of these published source assets took:
72
 
73
  | Bundle | First observed preparation with specialization | Subsequent cached preparation |
74
  | --- | ---: | ---: |
75
  | fast | 33 s | 0.060 s |
76
- | quality | 10 min 7 s | 0.037 s |
77
 
78
- Preparation includes model loading and compilation requested by Core AI, and
79
- excludes warmup and transcription. These are observed cache histories, not
80
- guaranteed fresh-install times: the underlying ANE cache state is not fully
81
- observable, and OS/device/application changes can require specialization again.
 
82
 
83
- AoT compiled artifacts
84
- are omitted because no significant load-time benefit has been demonstrated.
85
- Authoring debug locations were removed while preserving graph signatures and
86
- operation counts. `SHA256.json` lists distributed payload checksums.
 
14
  Choose one self-contained bundle:
15
 
16
  - `fast/`: full finite Fourier relative attention, 192 encoder positions,
17
+ fixed 15-second chunks, and an FP16 ANE decoder batching up to 128 independent chunks.
18
+ - `quality/`: full sinusoidal relative attention, shared weights for 576 and
19
+ 3,008 encoder positions (45-second and four-minute inputs), exact tiled
20
+ convolution subsampling, fixed 240-second chunks, and an FP32 CPU decoder
21
+ batching up to four independent chunks. Select the smallest shape that fits
22
+ each chunk; short recordings use the 45-second shape.
23
+
24
+ Both use W8A16 encoder weights with selected sensitive projections retained in
25
  FP16, and four concurrent encoder requests. Neither truncates attention within
26
+ its chunk or drops input audio. `metadata.json` supplies available input shapes;
27
+ `runtime.json` records the qualified scheduling policy. Weights are shared across
28
+ static functions in each asset. The labels describe a measured speed/context
29
+ tradeoff, not a universal quality ranking.
30
 
31
  Requires physical Apple silicon on macOS 27 or iOS 27 and a model-specific host
32
  runtime. Graphs accept model features and recurrent states, not audio files.
33
  The host must implement the frontend, greedy TDT loop and tokenizer described
34
+ by the metadata and sidecars. Vocabulary: 1,024 nonblank tokens, blank ID 1,024,
35
+ durations 0/1/2/3/4, and at most ten symbols per frame. Recurrent state is independent
36
+ between chunks. TDT emission frames and predicted durations support token/word
37
+ timings at an 80 ms frame step; they are native alignments, not forced alignment.
 
 
 
 
38
 
39
  ## Measured performance
40
 
41
  M3 MacBook Air (16 GB), macOS 27 build 26A428. Medians after warmup;
42
+ frontend, encoder and decoding are included. Preparation, file I/O and chunk
43
+ planning are excluded. Audio is JFK's **“We choose to go to the Moon”** speech:
44
+ a **20-second excerpt** and the **18-minute 15-second recording**.
 
45
 
46
  | Bundle | Audio duration | Transcription | Audio / elapsed |
47
  | --- | ---: | ---: | ---: |
48
  | fast | 20 seconds | 0.150 s | 133.4× |
49
  | fast | 18 min 15 s | 1.876 s | 584.0× |
50
+ | quality | 20 seconds | 0.153 s | 130.9× |
51
+ | quality | 18 min 15 s | 6.257 s | 175.0× |
52
+
53
+ Quality measurements use three timed runs after warmup, with stable tokens and
54
+ native word timings and nominal thermal state. The full recording uses 74 fast
55
+ chunks or five quality chunks. These contexts and execution policies differ.
56
 
57
  Whisper-normalized WER against the supplied long-recording reference:
58
 
59
+ | Bundle / reference | Word errors | WER |
60
  | --- | ---: | ---: |
61
  | fast | 79/2220 | 3.56% |
62
+ | quality | 45/2220 | 2.03% |
63
+ | FP32 source, same four-minute cuts as quality | 46/2220 | 2.07% |
64
+
65
+ With the benchmark's simpler normalization, quality scores 54/2219 (2.43%)
66
+ and the matching source 53/2219 (2.39%). This is one English recording, not a
67
+ general accuracy ranking. Quantization and floating-point operation order can
68
+ change token decisions. Both quality shapes pass the short-fixture projected
69
+ encoder check at 2.04% relative RMS versus the independent FP32 source and
70
+ produce identical valid outputs to each other. Word timings are ordered,
71
+ bounded and text-preserving; human word-boundary accuracy was not measured.
72
+
73
+ A bounded full-recording quality trace contains ANE predictions within all ten
74
+ encoder/subsampling calls and zero target GPU intervals. The decoder uses the
75
+ CPU. Compiler manifests mark both functions in each quality asset fully placed
76
+ on ANE; this is placement evidence, not an arithmetic-utilization measurement.
77
+ Sampled peak client-plus-attributed-neural memory is approximately 1.07 GB,
78
+ excluding unattributed compiler, driver and system memory. Phone execution of
79
+ this quality bundle has not been qualified.
80
 
81
  ## Device specialization and loading
82
 
83
+ Source `.aimodel` files specialize on the device. On the same Mac:
 
84
 
85
  | Bundle | First observed preparation with specialization | Subsequent cached preparation |
86
  | --- | ---: | ---: |
87
  | fast | 33 s | 0.060 s |
88
+ | quality | 6 min 2 s | 0.041 s |
89
 
90
+ Preparation includes Core AI model initialization and function loading;
91
+ it excludes host sidecar reads, audio I/O, chunk planning, warmup and transcription.
92
+ These are observed cache histories, not guaranteed fresh-install times. The
93
+ underlying ANE cache state is not fully observable, and OS/device/application
94
+ changes can require specialization again.
95
 
96
+ AoT compiled artifacts are omitted because no significant load-time benefit has
97
+ been demonstrated. Authoring debug locations were removed while preserving graph
98
+ signatures and operation counts. `SHA256.json` lists distributed payload checksums.
 
SHA256.json CHANGED
@@ -11,8 +11,8 @@
11
  },
12
  {
13
  "path": "README.md",
14
- "bytes": 4129,
15
- "sha256": "e49a5b9f1ea0fa4d4fa8a369f6f9b732da44aae3733b55ea19b696fc384911dc"
16
  },
17
  {
18
  "path": "fast/LICENSE",
@@ -137,7 +137,7 @@
137
  {
138
  "path": "quality/SHA256.json",
139
  "bytes": 2155,
140
- "sha256": "90a2d9ee273311455b6fb27425c8955cabc7190e0d0bb7cb3f83732d8de80ffb"
141
  },
142
  {
143
  "path": "quality/decoder.f32",
@@ -152,17 +152,17 @@
152
  {
153
  "path": "quality/encoder.aimodel/main.hash",
154
  "bytes": 32,
155
- "sha256": "1341561803e3f4ac7781b6dd61f8e1dbb366cd43bcd9ef7a45616f008d29bc5c"
156
  },
157
  {
158
  "path": "quality/encoder.aimodel/main.mlirb",
159
- "bytes": 673000984,
160
- "sha256": "6d1cf619f908597db03233b283cdee7bc1f0b91af650ffcbb98910a0d7b8e944"
161
  },
162
  {
163
  "path": "quality/encoder.aimodel/metadata.json",
164
  "bytes": 105,
165
- "sha256": "e3632aa1791fdec612c30f4e1c7416d81f346298ea3c3461f79f3a5c79fcec30"
166
  },
167
  {
168
  "path": "quality/mel-filter.f32",
@@ -171,28 +171,28 @@
171
  },
172
  {
173
  "path": "quality/metadata.json",
174
- "bytes": 648,
175
- "sha256": "4815b098209eb8e4d5fa5ba882d07da647c809af02e228092ce0fefdcf9f795a"
176
  },
177
  {
178
  "path": "quality/runtime.json",
179
- "bytes": 222,
180
- "sha256": "d0dadc946960eab57e7e2a18777d82de644540ea891737de687f1a48061be01d"
181
  },
182
  {
183
  "path": "quality/subsampling.aimodel/main.hash",
184
  "bytes": 32,
185
- "sha256": "0f26d1b91ad6bb0f1d391a6d451ce65189ee0b462181722323c3ffde20b36958"
186
  },
187
  {
188
  "path": "quality/subsampling.aimodel/main.mlirb",
189
- "bytes": 4403678,
190
- "sha256": "015ca4f66c9f65cd37cf201905f4c4cae7488c2338525327f48f60d294e3fc63"
191
  },
192
  {
193
  "path": "quality/subsampling.aimodel/metadata.json",
194
  "bytes": 105,
195
- "sha256": "fb03755e71f71e9b5410a65cce9f492309cd23fcf36d795d55db3f473512cd6c"
196
  },
197
  {
198
  "path": "quality/vocabulary.json",
 
11
  },
12
  {
13
  "path": "README.md",
14
+ "bytes": 5082,
15
+ "sha256": "cf41405ba0c95f42de280675abf725280ad7a682297b3f6d295b839bdb389214"
16
  },
17
  {
18
  "path": "fast/LICENSE",
 
137
  {
138
  "path": "quality/SHA256.json",
139
  "bytes": 2155,
140
+ "sha256": "3a04a0d4d3ff6defb0110aba0e98d678604424bb92b2b18539d06164073eefc3"
141
  },
142
  {
143
  "path": "quality/decoder.f32",
 
152
  {
153
  "path": "quality/encoder.aimodel/main.hash",
154
  "bytes": 32,
155
+ "sha256": "ce4c33f50678b578693d3e243d5cfb17015eca0478207ca115f3cb37bef22bd4"
156
  },
157
  {
158
  "path": "quality/encoder.aimodel/main.mlirb",
159
+ "bytes": 672997665,
160
+ "sha256": "80c2eba65224c1ea94e540310b4150bb9b5f90fc173d22d8e8fbdf46cda54131"
161
  },
162
  {
163
  "path": "quality/encoder.aimodel/metadata.json",
164
  "bytes": 105,
165
+ "sha256": "bab646772ac556e2994b81d16524116c979df6136b6db03e46024358c11c0a1c"
166
  },
167
  {
168
  "path": "quality/mel-filter.f32",
 
171
  },
172
  {
173
  "path": "quality/metadata.json",
174
+ "bytes": 746,
175
+ "sha256": "ca0164e863a5156c3123d22ef245110c0f4f538482acea013441aee199f0d8d3"
176
  },
177
  {
178
  "path": "quality/runtime.json",
179
+ "bytes": 224,
180
+ "sha256": "ce34a23fe78bb828a7c1a2db83a1bac101e8ef0858fee834d218e54a85ceddb8"
181
  },
182
  {
183
  "path": "quality/subsampling.aimodel/main.hash",
184
  "bytes": 32,
185
+ "sha256": "c81c114a265c14e125c95100f939126c5ff31c14cb91d2f04053a1aaa1f968f1"
186
  },
187
  {
188
  "path": "quality/subsampling.aimodel/main.mlirb",
189
+ "bytes": 4404378,
190
+ "sha256": "c36a9e7c2dcc8cb2ecae4675702c7d2a6755a8b18609fec43a9c8905e602a741"
191
  },
192
  {
193
  "path": "quality/subsampling.aimodel/metadata.json",
194
  "bytes": 105,
195
+ "sha256": "d27a50da08410c52dc818da0189fee6212906dd78d52732d14a24aca68a5d2f5"
196
  },
197
  {
198
  "path": "quality/vocabulary.json",
quality/SHA256.json CHANGED
@@ -22,17 +22,17 @@
22
  {
23
  "path": "encoder.aimodel/main.hash",
24
  "bytes": 32,
25
- "sha256": "1341561803e3f4ac7781b6dd61f8e1dbb366cd43bcd9ef7a45616f008d29bc5c"
26
  },
27
  {
28
  "path": "encoder.aimodel/main.mlirb",
29
- "bytes": 673000984,
30
- "sha256": "6d1cf619f908597db03233b283cdee7bc1f0b91af650ffcbb98910a0d7b8e944"
31
  },
32
  {
33
  "path": "encoder.aimodel/metadata.json",
34
  "bytes": 105,
35
- "sha256": "e3632aa1791fdec612c30f4e1c7416d81f346298ea3c3461f79f3a5c79fcec30"
36
  },
37
  {
38
  "path": "mel-filter.f32",
@@ -41,28 +41,28 @@
41
  },
42
  {
43
  "path": "metadata.json",
44
- "bytes": 648,
45
- "sha256": "4815b098209eb8e4d5fa5ba882d07da647c809af02e228092ce0fefdcf9f795a"
46
  },
47
  {
48
  "path": "runtime.json",
49
- "bytes": 222,
50
- "sha256": "d0dadc946960eab57e7e2a18777d82de644540ea891737de687f1a48061be01d"
51
  },
52
  {
53
  "path": "subsampling.aimodel/main.hash",
54
  "bytes": 32,
55
- "sha256": "0f26d1b91ad6bb0f1d391a6d451ce65189ee0b462181722323c3ffde20b36958"
56
  },
57
  {
58
  "path": "subsampling.aimodel/main.mlirb",
59
- "bytes": 4403678,
60
- "sha256": "015ca4f66c9f65cd37cf201905f4c4cae7488c2338525327f48f60d294e3fc63"
61
  },
62
  {
63
  "path": "subsampling.aimodel/metadata.json",
64
  "bytes": 105,
65
- "sha256": "fb03755e71f71e9b5410a65cce9f492309cd23fcf36d795d55db3f473512cd6c"
66
  },
67
  {
68
  "path": "vocabulary.json",
 
22
  {
23
  "path": "encoder.aimodel/main.hash",
24
  "bytes": 32,
25
+ "sha256": "ce4c33f50678b578693d3e243d5cfb17015eca0478207ca115f3cb37bef22bd4"
26
  },
27
  {
28
  "path": "encoder.aimodel/main.mlirb",
29
+ "bytes": 672997665,
30
+ "sha256": "80c2eba65224c1ea94e540310b4150bb9b5f90fc173d22d8e8fbdf46cda54131"
31
  },
32
  {
33
  "path": "encoder.aimodel/metadata.json",
34
  "bytes": 105,
35
+ "sha256": "bab646772ac556e2994b81d16524116c979df6136b6db03e46024358c11c0a1c"
36
  },
37
  {
38
  "path": "mel-filter.f32",
 
41
  },
42
  {
43
  "path": "metadata.json",
44
+ "bytes": 746,
45
+ "sha256": "ca0164e863a5156c3123d22ef245110c0f4f538482acea013441aee199f0d8d3"
46
  },
47
  {
48
  "path": "runtime.json",
49
+ "bytes": 224,
50
+ "sha256": "ce34a23fe78bb828a7c1a2db83a1bac101e8ef0858fee834d218e54a85ceddb8"
51
  },
52
  {
53
  "path": "subsampling.aimodel/main.hash",
54
  "bytes": 32,
55
+ "sha256": "c81c114a265c14e125c95100f939126c5ff31c14cb91d2f04053a1aaa1f968f1"
56
  },
57
  {
58
  "path": "subsampling.aimodel/main.mlirb",
59
+ "bytes": 4404378,
60
+ "sha256": "c36a9e7c2dcc8cb2ecae4675702c7d2a6755a8b18609fec43a9c8905e602a741"
61
  },
62
  {
63
  "path": "subsampling.aimodel/metadata.json",
64
  "bytes": 105,
65
+ "sha256": "d27a50da08410c52dc818da0189fee6212906dd78d52732d14a24aca68a5d2f5"
66
  },
67
  {
68
  "path": "vocabulary.json",
quality/encoder.aimodel/main.hash CHANGED
@@ -1 +1 @@
1
- m��Y}�23����{���P�˹��׸�D
 
1
+ ���R$���@1 AP��_��="����FͥA1
quality/encoder.aimodel/main.mlirb CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:6d1cf619f908597db03233b283cdee7bc1f0b91af650ffcbb98910a0d7b8e944
3
- size 673000984
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:80c2eba65224c1ea94e540310b4150bb9b5f90fc173d22d8e8fbdf46cda54131
3
+ size 672997665
quality/encoder.aimodel/metadata.json CHANGED
@@ -1,5 +1,5 @@
1
  {
2
- "creationDate" : "20260926T185446Z",
3
- "producer" : "coreai-core 1.0.0b2",
4
- "assetVersion" : "2.0"
5
  }
 
1
  {
2
+ "creationDate" : "20260927T202706Z",
3
+ "assetVersion" : "2.0",
4
+ "producer" : "coreai-core 1.0.0b2"
5
  }
quality/metadata.json CHANGED
@@ -3,7 +3,8 @@
3
  "repository": "nvidia/parakeet-tdt-0.6b-v2",
4
  "revision": "ae9ad07059c7c739ffaf932226a8fe64ae2620b0",
5
  "frame_buckets": [
6
- 3776
 
7
  ],
8
  "precision": "w8a16",
9
  "blank_id": 1024,
@@ -24,6 +25,9 @@
24
  "wide_spatial_projections": false,
25
  "fused_qkv": false,
26
  "split_glu": true,
 
 
 
27
  "subsampling_strategy": "tiled-exact",
28
  "subsampling_tile_frames": 192
29
  }
 
3
  "repository": "nvidia/parakeet-tdt-0.6b-v2",
4
  "revision": "ae9ad07059c7c739ffaf932226a8fe64ae2620b0",
5
  "frame_buckets": [
6
+ 576,
7
+ 3008
8
  ],
9
  "precision": "w8a16",
10
  "blank_id": 1024,
 
25
  "wide_spatial_projections": false,
26
  "fused_qkv": false,
27
  "split_glu": true,
28
+ "native_layer_norm": false,
29
+ "bounded_norm_variance": false,
30
+ "accurate_silu": true,
31
  "subsampling_strategy": "tiled-exact",
32
  "subsampling_tile_frames": 192
33
  }
quality/runtime.json CHANGED
@@ -1,7 +1,7 @@
1
  {
2
  "profile": "quality",
3
  "chunk_policy": "fixed",
4
- "chunk_seconds": 300,
5
  "encoder_concurrency": 4,
6
  "decoder": "cpu",
7
  "decoder_batch_size": 4,
 
1
  {
2
  "profile": "quality",
3
  "chunk_policy": "fixed",
4
+ "chunk_seconds": 240.0,
5
  "encoder_concurrency": 4,
6
  "decoder": "cpu",
7
  "decoder_batch_size": 4,
quality/subsampling.aimodel/main.hash CHANGED
@@ -1 +1 @@
1
- \��l�e�7� ����H�#8RS'�`Ҕ��c
 
1
+ �j�|-̌��Fup,}*gU��� ��:����A
quality/subsampling.aimodel/main.mlirb CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:015ca4f66c9f65cd37cf201905f4c4cae7488c2338525327f48f60d294e3fc63
3
- size 4403678
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c36a9e7c2dcc8cb2ecae4675702c7d2a6755a8b18609fec43a9c8905e602a741
3
+ size 4404378
quality/subsampling.aimodel/metadata.json CHANGED
@@ -1,5 +1,5 @@
1
  {
2
- "producer" : "coreai-core 1.0.0b2",
3
  "assetVersion" : "2.0",
4
- "creationDate" : "20260926T185448Z"
5
  }
 
1
  {
2
+ "creationDate" : "20260927T202709Z",
3
  "assetVersion" : "2.0",
4
+ "producer" : "coreai-core 1.0.0b2"
5
  }