File size: 4,499 Bytes
8f2d832
 
 
 
 
7b8d7a2
8f2d832
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
573da8e
8f2d832
 
 
 
 
 
 
 
 
 
573da8e
8f2d832
573da8e
 
8f2d832
 
 
 
 
 
 
 
 
9de0fa5
de76e2e
8f2d832
 
de76e2e
 
 
8f2d832
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9de0fa5
8f2d832
012b42d
 
 
 
 
 
 
 
 
 
 
 
a4436fe
012b42d
 
 
 
8f2d832
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
---
license: apache-2.0
library_name: coreml
pipeline_tag: text-classification
base_model: convaiinnovations/laya
base_model_relation: quantized
tags:
  - coreml
  - laya
  - apple-silicon
  - neural-engine
  - decision-model
  - fluidaudio
---

# laya-coreml

Core ML conversion of **laya-multilingual** (Convai Innovations, Apache-2.0): a 322M-parameter
mmBERT-base encoder with a typed decision head that answers `choice`, `score`, and `noul`
questions about a text state in one forward pass, returning calibrated probabilities and no
generated tokens. Weights are unchanged from
[`convaiinnovations/laya`](https://huggingface.co/convaiinnovations/laya) `multilingual/` at
revision `1c5edc17a7acd8701df6fc341c0d179f1c62c982`.

Runs through [FluidUse](https://github.com/FluidInference/FluidUse) (`LayaManager`) on macOS 14+.

```swift
let laya = try await LayaManager.load()  // downloads the 128 + 512 buckets and tokenizer.json
let answer = try await laya.answer(
    state: "The T piece dropped at column 3 leaves one hole under it.",
    question: .noul("Is this a clean placement?"))
print(answer.noul!)  // P(true)
```

```bash
swift run -c release FluidUseLaya answer --state "…" --type choice \
    --instructions "What does the customer want?" --options "refund|order status|technical help"
swift run -c release FluidUseLaya tetris      # headless Tetris played by laya decisions
swift run -c release LayaTetrisDemo           # SwiftUI demo
```

## Files

| File | Tokens | Notes |
| --- | ---: | --- |
| `laya_multilingual_fp16_L128_options32.mlmodelc` | 128 | Short prompts; runs on CPU + Neural Engine |
| `laya_multilingual_fp16_L256_options32.mlmodelc` | 256 | |
| `laya_multilingual_fp16_L512_options32.mlmodelc` | 512 | Long states; GPU is faster than ANE here |
| `laya_multilingual_fp16_L1024_options32.mlmodelc` | 1024 | Upstream `max_len`; GPU |
| `laya_multilingual_e8_L{128,256,512,1024}_options32.mlmodelc` | | Same buckets with an int8 embedding table: 448–453 MB each, accuracy within 0.5 points of fp16 on the full benchmark |
| `tokenizer.json` | | mmBERT / Gemma vocabulary (256k), byte fallback |

Each fp16 bucket is a complete model (614 MB, 393 MB of which is the embedding table) with
32 option slots; the `e8` buckets store that table as int8 per-channel. Encoder-weight int8 and
6-/4-bit palettes fail the parity gates (the ANE in particular), so they are not published. `FluidUse` picks the smallest loaded bucket that fits a prompt and truncates
the state on the right for the largest one, exactly like laya's `max_len`.

Inputs: `input_ids` int32 `[1, L]`, `attention_mask` int32 `[1, L]`, `marker_map` float32
`[1, 32, L]` (one-hot `[MASK]` position per option), `question_type` float32 `[1, 3]`.
Outputs: `logits` `[1, 32]`, `probabilities` `[1, 32]`, `action_probabilities` `[1, 2]`.
Sequence format: `[CLS] <type> question: <instructions> [SEP] ([MASK] <option>)* [SEP] <state> [SEP]`.

## Parity and latency

Apple M5 Pro, macOS 27.0, 16 fixture questions vs. the unmodified PyTorch FP32 runtime:
16/16 argmax agreement on every bucket and compute-unit setting, max probability error 0.0021
(`ALL`) / 0.0126 (`CPU_AND_NE`). Per-question latency, warm:

| Bucket | CPU + ANE | All units |
| --- | ---: | ---: |
| L128 | **3.6 ms** | 3.9 ms |
| L256 | 9.9 ms | **5.2 ms** |
| L512 | 27.5 ms | **9.0 ms** |
| L1024 | 80.1 ms | **17.9 ms** |

On laya's published application suites (3,899 questions, seed 13, rebuilt from upstream's scripts),
the Core ML buckets answered from Swift match the PyTorch reference's accuracy on every suite at
5.2 ms median per question (p95 18 ms):

| Suite | Upstream (T4, PyTorch) | Core ML (M5 Pro) |
| --- | ---: | ---: |
| jev.ag_news | 0.930 | **0.935** |
| jev.emotion | 0.530 | **0.537** |
| massive_intent.en | 0.657 | **0.657** |
| app.support_triage | 0.522 | **0.542** |
| app.email_spam | 0.993 | **0.993** |
| app.phishing | 0.993 | **0.993** |
| app.guardrails_jailbreak | 0.755 | **0.808** |
| app.moderation_toxicity | 0.525 | **0.535** |
| app.rag_relevance | 0.657 | **0.672** |
| app.model_routing_domain | 0.123 | **0.441** |

Conversion pipeline, verification reports, and Swift parity fixtures:
[mobius `models/computer-use/laya/coreml`](https://github.com/FluidInference/mobius).

## License

Apache-2.0, following the upstream weights and code by Convai Innovations
([NandhaKishorM/laya](https://github.com/NandhaKishorM/laya)). Independent conversion; not an
official Convai Innovations release.