mlboydaisuke commited on
Commit
7964797
·
verified ·
1 Parent(s): 820d4e9

README: ios/ is the JIT .aimodel, ios-h19p/ the h19p bundle (moved in 820d4e90); SHA256SUMS for the new layout

Browse files
Files changed (2) hide show
  1. README.md +24 -16
  2. SHA256SUMS +28 -16
README.md CHANGED
@@ -37,8 +37,8 @@ There is no training for your label set and no generated text; every label gets
37
  DeBERTa-v3-large encoder with a trained label head, 486M parameters with its 128k-token embedding
38
  table (340M on Fastino's card). Its classification path is one static Core AI graph in fp16; the
39
  tokenizer, the schema layout and the softmax / sigmoid run in the host. One call takes 37.8 ms on an
40
- iPhone 18 Pro (256-token graph, thermal state nominal), and its decisions equal the fp32 `gliner2`
41
- library's on every one of 787 decisions over 454 texts.
42
 
43
  ## Use it
44
 
@@ -112,13 +112,15 @@ decisions.
112
  | fp16 graph S = 256, Mac GPU (M4 Max, macOS 27.0) | 361 / 606 | 606 | 0.014 | 0.0022 |
113
  | fp16 graph S = 512, Mac GPU | 454 / 787 | 787 | 0.018 | 0.0025 |
114
  | Swift host (CoreAIKit `TextClassifier`), Mac GPU | 454 / 787 | 787 | 0.014 | — |
115
- | fp16 S = 256, iPhone 18 Pro GPU (h19p) | 361 / 606 | 606 | 0.019 | 0.0023 |
116
- | fp16 S = 512, iPhone 18 Pro GPU (h19p) | 454 / 787 | 787 | 0.019 | 0.0023 |
117
 
118
  The Swift host's token ids equal `gliner2`'s on all 454 texts, and its Mac GPU logits equal the Python
119
  engine run's bit for bit (5,749 of 5,749 values). The iPhone's logits are within 0.016 of the Mac
120
  GPU's, with every decision the same. iPhone rows: iOS 27.0 (build 24A437), measured 2026-09-26 with the
121
- zoo's gate app.
 
 
122
 
123
  On the fixture's 340 short rows, one call per row with all of its heads and the default threshold, 63.5
124
  % of the decisions equal the dataset's gold label. That is this port's number on the development
@@ -131,11 +133,14 @@ One call, warm, median over 100 calls after 5 warm-up calls on fixture inputs:
131
 
132
  | device | S = 256 | S = 512 | load, second time | footprint after load |
133
  |---|---|---|---|---|
134
- | iPhone 18 Pro GPU, thermal state nominal, phone rested 7 min | 37.8 ms (p90 38.1) | 92.9 ms (p90 94.5) | 0.14 s / 0.84 s | 220 MB / 376 MB |
135
  | M4 Max GPU, another job on the GPU | 28 ms | 52 ms | 0.01 s | — |
136
 
137
- The first load after installing on the iPhone took 1.3 s (S = 256) and 1.8 s (S = 512), with a first
138
- call of 1.2 s and 0.4 s. Several hundred calls without a pause slow the iPhone: across one loop of the
 
 
 
139
  S = 512 fixture a call went from 92 to 151 ms, and a run started on a warm phone measured 84 ms at
140
  S = 256 against 38 ms rested. Five minutes of rest brings the speed back.
141
 
@@ -145,19 +150,22 @@ S = 256 against 38 ms rested. Five minutes of rest brings the speed back.
145
  |---|---|---|
146
  | `macos/gliner25-decide_float16_s256_m32.aimodel` | JIT bundle, S = 256 | 873 MB |
147
  | `macos/gliner25-decide_float16_s512_m32.aimodel` | JIT bundle, S = 512 | 875 MB |
148
- | `ios/gliner25-decide_float16_s256_m32.h19p.aimodelc` | compiled ahead of time for the iPhone 18 Pro GPU (h19p) | 974 MB |
149
- | `ios/gliner25-decide_float16_s512_m32.h19p.aimodelc` | same, S = 512 | 976 MB |
150
- | `macos/tokenizer/`, `ios/tokenizer/` | the DeBERTa-v3 SentencePiece tokenizer, declared as `XLMRobertaTokenizer` so swift-transformers loads it | 8.3 MB |
151
- | `macos/classifier.json`, `ios/classifier.json` | the graph contract, the two shapes, the marker token ids, the host rules | |
152
- | `macos/reference_s*.json`, `ios/reference_s*.json` | one fixture text with its graph inputs and fp32 logits, for a host to check itself against | |
 
153
  | `gate/` | the fixture: the 21 card examples and the fast-decisions rows, each with its token ids, marker positions, fp32 logits and decision | 5 MB |
154
  | `LICENSE`, `NOTICE`, `source/` | Apache-2.0, the origin and what was converted, the source `config.json` files | |
155
  | `config.json` | marks the repo as Core AI `.aimodel` bundles for the zoo's tooling | |
156
  | `SHA256SUMS` | every file's checksum; `conversion/gliner25_decide/stage_ship.py --check <dir>` in the zoo verifies a download | |
157
 
158
- The iPhone bundles load only on the h19p architecture (iPhone 18 Pro). An h18p bundle for the iPhone
159
- 17 Pro compiles from the same `.aimodel` with the zoo recipe, but has not been run on that device and
160
- is not shipped.
 
 
161
 
162
  ## Limits
163
 
 
37
  DeBERTa-v3-large encoder with a trained label head, 486M parameters with its 128k-token embedding
38
  table (340M on Fastino's card). Its classification path is one static Core AI graph in fp16; the
39
  tokenizer, the schema layout and the softmax / sigmoid run in the host. One call takes 37.8 ms on an
40
+ iPhone 18 Pro (256-token graph compiled ahead of time for it, thermal state nominal), and its decisions
41
+ equal the fp32 `gliner2` library's on every one of 787 decisions over 454 texts.
42
 
43
  ## Use it
44
 
 
112
  | fp16 graph S = 256, Mac GPU (M4 Max, macOS 27.0) | 361 / 606 | 606 | 0.014 | 0.0022 |
113
  | fp16 graph S = 512, Mac GPU | 454 / 787 | 787 | 0.018 | 0.0025 |
114
  | Swift host (CoreAIKit `TextClassifier`), Mac GPU | 454 / 787 | 787 | 0.014 | — |
115
+ | fp16 S = 256, iPhone 18 Pro GPU, AOT (h19p) | 361 / 606 | 606 | 0.019 | 0.0023 |
116
+ | fp16 S = 512, iPhone 18 Pro GPU, AOT (h19p) | 454 / 787 | 787 | 0.019 | 0.0023 |
117
 
118
  The Swift host's token ids equal `gliner2`'s on all 454 texts, and its Mac GPU logits equal the Python
119
  engine run's bit for bit (5,749 of 5,749 values). The iPhone's logits are within 0.016 of the Mac
120
  GPU's, with every decision the same. iPhone rows: iOS 27.0 (build 24A437), measured 2026-09-26 with the
121
+ zoo's gate app. On the same phone, the JIT bundles now in `ios/` matched the reference on every decision
122
+ (606 of 606 at S = 256, 787 of 787 at S = 512), with logits within 0.016 of it
123
+ ([knowledge/gliner25-decide.md](https://github.com/john-rocky/coreai-model-zoo/blob/main/knowledge/gliner25-decide.md) §7).
124
 
125
  On the fixture's 340 short rows, one call per row with all of its heads and the default threshold, 63.5
126
  % of the decisions equal the dataset's gold label. That is this port's number on the development
 
133
 
134
  | device | S = 256 | S = 512 | load, second time | footprint after load |
135
  |---|---|---|---|---|
136
+ | iPhone 18 Pro GPU, AOT (h19p), thermal state nominal, phone rested 7 min | 37.8 ms (p90 38.1) | 92.9 ms (p90 94.5) | 0.14 s / 0.84 s | 220 MB / 376 MB |
137
  | M4 Max GPU, another job on the GPU | 28 ms | 52 ms | 0.01 s | — |
138
 
139
+ The first load of the AOT bundles after installing on the iPhone took 1.3 s (S = 256) and 1.8 s
140
+ (S = 512), with a first call of 1.2 s and 0.4 s. The JIT bundles now in `ios/`, on the same phone
141
+ (rested, thermal state nominal): first load after installing 1.48 s and 2.32 s, first call 1.39 s and
142
+ 0.49 s, load after a relaunch 0.36 s and 0.13 s, one call 35.9 ms and 87.2 ms (median;
143
+ [knowledge/gliner25-decide.md](https://github.com/john-rocky/coreai-model-zoo/blob/main/knowledge/gliner25-decide.md) §7). Several hundred calls without a pause slow the iPhone: across one loop of the
144
  S = 512 fixture a call went from 92 to 151 ms, and a run started on a warm phone measured 84 ms at
145
  S = 256 against 38 ms rested. Five minutes of rest brings the speed back.
146
 
 
150
  |---|---|---|
151
  | `macos/gliner25-decide_float16_s256_m32.aimodel` | JIT bundle, S = 256 | 873 MB |
152
  | `macos/gliner25-decide_float16_s512_m32.aimodel` | JIT bundle, S = 512 | 875 MB |
153
+ | `ios/gliner25-decide_float16_s256_m32.aimodel`, `ios/gliner25-decide_float16_s512_m32.aimodel` | the same two JIT bundles, byte for byte; an iPhone specializes them on its first load | 873 MB, 875 MB |
154
+ | `ios-h19p/gliner25-decide_float16_s256_m32.h19p.aimodelc` | compiled ahead of time for the iPhone 18 Pro GPU (h19p); moved from `ios/` in revision `820d4e90` (2026-09-26) | 974 MB |
155
+ | `ios-h19p/gliner25-decide_float16_s512_m32.h19p.aimodelc` | same, S = 512 | 976 MB |
156
+ | `macos/tokenizer/`, `ios/tokenizer/`, `ios-h19p/tokenizer/` | the DeBERTa-v3 SentencePiece tokenizer, declared as `XLMRobertaTokenizer` so swift-transformers loads it | 8.3 MB |
157
+ | `macos/classifier.json`, `ios/classifier.json`, `ios-h19p/classifier.json` | the graph contract, the two shapes, the marker token ids, the host rules; `ios-h19p/`'s names the `.h19p.aimodelc` bundles | |
158
+ | `macos/reference_s*.json`, `ios/reference_s*.json`, `ios-h19p/reference_s*.json` | one fixture text with its graph inputs and fp32 logits, for a host to check itself against | |
159
  | `gate/` | the fixture: the 21 card examples and the fast-decisions rows, each with its token ids, marker positions, fp32 logits and decision | 5 MB |
160
  | `LICENSE`, `NOTICE`, `source/` | Apache-2.0, the origin and what was converted, the source `config.json` files | |
161
  | `config.json` | marks the repo as Core AI `.aimodel` bundles for the zoo's tooling | |
162
  | `SHA256SUMS` | every file's checksum; `conversion/gliner25_decide/stage_ship.py --check <dir>` in the zoo verifies a download | |
163
 
164
+ The JIT bundles in `ios/` have been run on the iPhone 18 Pro only (numbers above). The `ios-h19p/`
165
+ bundles load only on the h19p architecture (iPhone 18 Pro): the runtime refuses a compiled bundle on
166
+ another architecture ([knowledge/jit-distribution.md](https://github.com/john-rocky/coreai-model-zoo/blob/main/knowledge/jit-distribution.md)). An h18p bundle for the iPhone
167
+ 17 Pro compiles from the same `.aimodel` with the zoo recipe, but has not been run on that device and is
168
+ not shipped.
169
 
170
  ## Limits
171
 
SHA256SUMS CHANGED
@@ -1,25 +1,37 @@
1
  c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4 LICENSE
2
  b1a4310c8287acc48e243059357a0936627d2a079e6783d8a92f9011c445be4a NOTICE
3
- 1d96f6a1b558d446dda416cc2c805a0b5100370be4ebc237c9923e67921bced9 README.md
4
  8bb12dab02fee45172391da555fb2ba75d2d341e69edb574e59412f2c801f6a7 config.json
5
  967c3d87734a3c939b8f73c66ae7b7c41cf3e5061191eeeba3d6f4a9729b26c2 gate/fast_decisions_long.json
6
  7b017250d6a2bfdd9952f799ebb777ba074d277052d9aa6fd486edfcb8e09f71 gate/fast_decisions_s256.json
7
  8b756164f153395805fc4c89b029352c441e25a6c4006c2b1572d7f427f53396 gate/readme21.json
8
- 4273ccd1358d96bbe8d43bae8c841ee108c8a2f05a43d76d2e2780bd5aca9c10 ios/classifier.json
9
- bf6cf50b4cae51e8cc6843c0bc4575c068281b80f06ca2bea42bc9c761258b0c ios/gliner25-decide_float16_s256_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/manifest.plist
10
- 7aaa044a031196db3a49a7664d6e2364785d224dd35634bb7dd70de7a652c6b2 ios/gliner25-decide_float16_s256_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/resources.bin
11
- 4c9199afa2abb0be4445889cc8b82bd3ed217c08846c905f5fe25e476c656952 ios/gliner25-decide_float16_s256_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/specialized_model_0.mpsgraph
12
- 6c2bde46dd704ae6111c788ea16cd2419b6ec772c480df5348360bc033958b5d ios/gliner25-decide_float16_s256_m32.h19p.aimodelc/main-h19p.mlirb
13
- fdb4c639a972bc089b723e8d8f5c3d062f065afb05ecb14675079953616fee5a ios/gliner25-decide_float16_s256_m32.h19p.aimodelc/main.hash
14
- 920ddc4e25d39ecb989b55809337859ecdf9ad192d4fa45e2ab490afd879194d ios/gliner25-decide_float16_s256_m32.h19p.aimodelc/metadata.json
15
- acc58821a1f8382351e4f031eb32c157fa70a063664344ca5e497015e23e109a ios/gliner25-decide_float16_s256_m32.h19p.aimodelc/stats.json
16
- 4890a364c505a4e9b73830add7981dc3f613ea8d837f21522e942bf122672715 ios/gliner25-decide_float16_s512_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/manifest.plist
17
- df0e15b3f643899312718d7d78b63d2a7428c0e73a0aa338f906e5612a862e51 ios/gliner25-decide_float16_s512_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/resources.bin
18
- 01b1195d27d96546d43c821fde3167489d6e25b70e9b03bd1eb6799c4e773afb ios/gliner25-decide_float16_s512_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/specialized_model_0.mpsgraph
19
- 02a775c052dcd1c6d10df805ba2a29667b5846af6d8f7d912a1b9e6b93f3cfde ios/gliner25-decide_float16_s512_m32.h19p.aimodelc/main-h19p.mlirb
20
- 4b78d31737fffbaa5917ef801ecd2b797a83331a02cf8acf12b5f856faeda0e8 ios/gliner25-decide_float16_s512_m32.h19p.aimodelc/main.hash
21
- bbda0a3f7d8d58b48dd4ee58b3a0089772e2931bba3826c63d28d7a66b6cae42 ios/gliner25-decide_float16_s512_m32.h19p.aimodelc/metadata.json
22
- 64a5dc9130abf04e4130e4288bb09cab09bc7b57b804c3b6cbb3b8cb2281fa9b ios/gliner25-decide_float16_s512_m32.h19p.aimodelc/stats.json
 
 
 
 
 
 
 
 
 
 
 
 
23
  6bf81a20cd9bec856d4cd6b6e8d0b33219d22a82b9f3a01dd26be7013390ac5e ios/reference_s256.json
24
  a46e93d521cd43494f4ec311245698546f14fc32428a2b02471fbe7ce3c00e66 ios/reference_s512.json
25
  84ea70143f533d7e99b393d87f20010887a9ac2cba955828ef313886e4e83f4f ios/tokenizer/special_tokens_map.json
 
1
  c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4 LICENSE
2
  b1a4310c8287acc48e243059357a0936627d2a079e6783d8a92f9011c445be4a NOTICE
3
+ 3a44a7ffb06ace964bab7948cf32c6e438d822bef24cf44e8580776fe48d6bc0 README.md
4
  8bb12dab02fee45172391da555fb2ba75d2d341e69edb574e59412f2c801f6a7 config.json
5
  967c3d87734a3c939b8f73c66ae7b7c41cf3e5061191eeeba3d6f4a9729b26c2 gate/fast_decisions_long.json
6
  7b017250d6a2bfdd9952f799ebb777ba074d277052d9aa6fd486edfcb8e09f71 gate/fast_decisions_s256.json
7
  8b756164f153395805fc4c89b029352c441e25a6c4006c2b1572d7f427f53396 gate/readme21.json
8
+ 4273ccd1358d96bbe8d43bae8c841ee108c8a2f05a43d76d2e2780bd5aca9c10 ios-h19p/classifier.json
9
+ bf6cf50b4cae51e8cc6843c0bc4575c068281b80f06ca2bea42bc9c761258b0c ios-h19p/gliner25-decide_float16_s256_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/manifest.plist
10
+ 7aaa044a031196db3a49a7664d6e2364785d224dd35634bb7dd70de7a652c6b2 ios-h19p/gliner25-decide_float16_s256_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/resources.bin
11
+ 4c9199afa2abb0be4445889cc8b82bd3ed217c08846c905f5fe25e476c656952 ios-h19p/gliner25-decide_float16_s256_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/specialized_model_0.mpsgraph
12
+ 6c2bde46dd704ae6111c788ea16cd2419b6ec772c480df5348360bc033958b5d ios-h19p/gliner25-decide_float16_s256_m32.h19p.aimodelc/main-h19p.mlirb
13
+ fdb4c639a972bc089b723e8d8f5c3d062f065afb05ecb14675079953616fee5a ios-h19p/gliner25-decide_float16_s256_m32.h19p.aimodelc/main.hash
14
+ 920ddc4e25d39ecb989b55809337859ecdf9ad192d4fa45e2ab490afd879194d ios-h19p/gliner25-decide_float16_s256_m32.h19p.aimodelc/metadata.json
15
+ acc58821a1f8382351e4f031eb32c157fa70a063664344ca5e497015e23e109a ios-h19p/gliner25-decide_float16_s256_m32.h19p.aimodelc/stats.json
16
+ 4890a364c505a4e9b73830add7981dc3f613ea8d837f21522e942bf122672715 ios-h19p/gliner25-decide_float16_s512_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/manifest.plist
17
+ df0e15b3f643899312718d7d78b63d2a7428c0e73a0aa338f906e5612a862e51 ios-h19p/gliner25-decide_float16_s512_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/resources.bin
18
+ 01b1195d27d96546d43c821fde3167489d6e25b70e9b03bd1eb6799c4e773afb ios-h19p/gliner25-decide_float16_s512_m32.h19p.aimodelc/main-h19p-delegates/MPSGraph/mpsExecutable.mpsgraphpackage/specialized_model_0.mpsgraph
19
+ 02a775c052dcd1c6d10df805ba2a29667b5846af6d8f7d912a1b9e6b93f3cfde ios-h19p/gliner25-decide_float16_s512_m32.h19p.aimodelc/main-h19p.mlirb
20
+ 4b78d31737fffbaa5917ef801ecd2b797a83331a02cf8acf12b5f856faeda0e8 ios-h19p/gliner25-decide_float16_s512_m32.h19p.aimodelc/main.hash
21
+ bbda0a3f7d8d58b48dd4ee58b3a0089772e2931bba3826c63d28d7a66b6cae42 ios-h19p/gliner25-decide_float16_s512_m32.h19p.aimodelc/metadata.json
22
+ 64a5dc9130abf04e4130e4288bb09cab09bc7b57b804c3b6cbb3b8cb2281fa9b ios-h19p/gliner25-decide_float16_s512_m32.h19p.aimodelc/stats.json
23
+ 6bf81a20cd9bec856d4cd6b6e8d0b33219d22a82b9f3a01dd26be7013390ac5e ios-h19p/reference_s256.json
24
+ a46e93d521cd43494f4ec311245698546f14fc32428a2b02471fbe7ce3c00e66 ios-h19p/reference_s512.json
25
+ 84ea70143f533d7e99b393d87f20010887a9ac2cba955828ef313886e4e83f4f ios-h19p/tokenizer/special_tokens_map.json
26
+ 3ad87d9ffe669147063e70850927dd2da90249e2acc5c8527f1eb65df467bcc8 ios-h19p/tokenizer/tokenizer.json
27
+ 07425ad6c0219f1d891fbc3e712737f949fd0dcfa3caef25c1a01aa2bf304c1d ios-h19p/tokenizer/tokenizer_config.json
28
+ 27046ec5a1c4932d0d6ff2b0ca60c3b3f17667d29049291d7ec5c80768ccbb81 ios/classifier.json
29
+ 1289dde3ae9e1d931f4afc1f42cb12fb129ce42fca901775a83d108aba06c38a ios/gliner25-decide_float16_s256_m32.aimodel/main.hash
30
+ 65dc8934552fa866bb70cd5e3d7c65b2b08a51b0d38f7c21098f5bec17aa940d ios/gliner25-decide_float16_s256_m32.aimodel/main.mlirb
31
+ fc01837eccd833a52f12386f11162afa4ce0a687349c9a8cd35f198a891c428c ios/gliner25-decide_float16_s256_m32.aimodel/metadata.json
32
+ 9c7a5996ececfef86de4564d18b14020fdf9657fdc51447d569482288e513fea ios/gliner25-decide_float16_s512_m32.aimodel/main.hash
33
+ b5e22f647c7de2bff26d2036341eaac0944e12be364efdbf7cc72434f45f563d ios/gliner25-decide_float16_s512_m32.aimodel/main.mlirb
34
+ 8c2c51bea2d37edebf9b9e365efb65a2903209b4965eb493cbf34e0f820c80a1 ios/gliner25-decide_float16_s512_m32.aimodel/metadata.json
35
  6bf81a20cd9bec856d4cd6b6e8d0b33219d22a82b9f3a01dd26be7013390ac5e ios/reference_s256.json
36
  a46e93d521cd43494f4ec311245698546f14fc32428a2b02471fbe7ce3c00e66 ios/reference_s512.json
37
  84ea70143f533d7e99b393d87f20010887a9ac2cba955828ef313886e4e83f4f ios/tokenizer/special_tokens_map.json