Model card: L1024
Browse files
README.md
CHANGED
|
@@ -45,6 +45,7 @@ swift run -c release fluidaudiocli laya-tetris # headless Tetris played by lay
|
|
| 45 |
| `laya_multilingual_fp16_L128_options32.mlmodelc` | 128 | Short prompts; runs on CPU + Neural Engine |
|
| 46 |
| `laya_multilingual_fp16_L256_options32.mlmodelc` | 256 | |
|
| 47 |
| `laya_multilingual_fp16_L512_options32.mlmodelc` | 512 | Long states; GPU is faster than ANE here |
|
|
|
|
| 48 |
| `tokenizer.json` | | mmBERT / Gemma vocabulary (256k), byte fallback |
|
| 49 |
|
| 50 |
Each bucket is a complete FP16 model (614 MB, 393 MB of which is the embedding table) with
|
|
@@ -67,6 +68,7 @@ Apple M5 Pro, macOS 27.0, 16 fixture questions vs. the unmodified PyTorch FP32 r
|
|
| 67 |
| L128 | **3.6 ms** | 3.9 ms |
|
| 68 |
| L256 | 9.9 ms | **5.2 ms** |
|
| 69 |
| L512 | 27.5 ms | **9.0 ms** |
|
|
|
|
| 70 |
|
| 71 |
Conversion pipeline, verification reports, and Swift parity fixtures:
|
| 72 |
[mobius `models/computer-use/laya/coreml`](https://github.com/FluidInference/mobius).
|
|
|
|
| 45 |
| `laya_multilingual_fp16_L128_options32.mlmodelc` | 128 | Short prompts; runs on CPU + Neural Engine |
|
| 46 |
| `laya_multilingual_fp16_L256_options32.mlmodelc` | 256 | |
|
| 47 |
| `laya_multilingual_fp16_L512_options32.mlmodelc` | 512 | Long states; GPU is faster than ANE here |
|
| 48 |
+
| `laya_multilingual_fp16_L1024_options32.mlmodelc` | 1024 | Upstream `max_len`; GPU |
|
| 49 |
| `tokenizer.json` | | mmBERT / Gemma vocabulary (256k), byte fallback |
|
| 50 |
|
| 51 |
Each bucket is a complete FP16 model (614 MB, 393 MB of which is the embedding table) with
|
|
|
|
| 68 |
| L128 | **3.6 ms** | 3.9 ms |
|
| 69 |
| L256 | 9.9 ms | **5.2 ms** |
|
| 70 |
| L512 | 27.5 ms | **9.0 ms** |
|
| 71 |
+
| L1024 | 80.1 ms | **17.9 ms** |
|
| 72 |
|
| 73 |
Conversion pipeline, verification reports, and Swift parity fixtures:
|
| 74 |
[mobius `models/computer-use/laya/coreml`](https://github.com/FluidInference/mobius).
|