alexwengg commited on
Commit
9de0fa5
·
verified ·
1 Parent(s): 1be4eb9

Model card: L1024

Browse files
Files changed (1) hide show
  1. README.md +2 -0
README.md CHANGED
@@ -45,6 +45,7 @@ swift run -c release fluidaudiocli laya-tetris # headless Tetris played by lay
45
  | `laya_multilingual_fp16_L128_options32.mlmodelc` | 128 | Short prompts; runs on CPU + Neural Engine |
46
  | `laya_multilingual_fp16_L256_options32.mlmodelc` | 256 | |
47
  | `laya_multilingual_fp16_L512_options32.mlmodelc` | 512 | Long states; GPU is faster than ANE here |
 
48
  | `tokenizer.json` | | mmBERT / Gemma vocabulary (256k), byte fallback |
49
 
50
  Each bucket is a complete FP16 model (614 MB, 393 MB of which is the embedding table) with
@@ -67,6 +68,7 @@ Apple M5 Pro, macOS 27.0, 16 fixture questions vs. the unmodified PyTorch FP32 r
67
  | L128 | **3.6 ms** | 3.9 ms |
68
  | L256 | 9.9 ms | **5.2 ms** |
69
  | L512 | 27.5 ms | **9.0 ms** |
 
70
 
71
  Conversion pipeline, verification reports, and Swift parity fixtures:
72
  [mobius `models/computer-use/laya/coreml`](https://github.com/FluidInference/mobius).
 
45
  | `laya_multilingual_fp16_L128_options32.mlmodelc` | 128 | Short prompts; runs on CPU + Neural Engine |
46
  | `laya_multilingual_fp16_L256_options32.mlmodelc` | 256 | |
47
  | `laya_multilingual_fp16_L512_options32.mlmodelc` | 512 | Long states; GPU is faster than ANE here |
48
+ | `laya_multilingual_fp16_L1024_options32.mlmodelc` | 1024 | Upstream `max_len`; GPU |
49
  | `tokenizer.json` | | mmBERT / Gemma vocabulary (256k), byte fallback |
50
 
51
  Each bucket is a complete FP16 model (614 MB, 393 MB of which is the embedding table) with
 
68
  | L128 | **3.6 ms** | 3.9 ms |
69
  | L256 | 9.9 ms | **5.2 ms** |
70
  | L512 | 27.5 ms | **9.0 ms** |
71
+ | L1024 | 80.1 ms | **17.9 ms** |
72
 
73
  Conversion pipeline, verification reports, and Swift parity fixtures:
74
  [mobius `models/computer-use/laya/coreml`](https://github.com/FluidInference/mobius).