Model card: benchmark table
Browse files
README.md
CHANGED
|
@@ -70,6 +70,23 @@ Apple M5 Pro, macOS 27.0, 16 fixture questions vs. the unmodified PyTorch FP32 r
|
|
| 70 |
| L512 | 27.5 ms | **9.0 ms** |
|
| 71 |
| L1024 | 80.1 ms | **17.9 ms** |
|
| 72 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 73 |
Conversion pipeline, verification reports, and Swift parity fixtures:
|
| 74 |
[mobius `models/computer-use/laya/coreml`](https://github.com/FluidInference/mobius).
|
| 75 |
|
|
|
|
| 70 |
| L512 | 27.5 ms | **9.0 ms** |
|
| 71 |
| L1024 | 80.1 ms | **17.9 ms** |
|
| 72 |
|
| 73 |
+
On laya's published application suites (3,899 questions, seed 13, rebuilt from upstream's scripts),
|
| 74 |
+
the Core ML buckets answered from Swift match the PyTorch reference's accuracy on every suite at
|
| 75 |
+
5.2 ms median per question (p95 18 ms):
|
| 76 |
+
|
| 77 |
+
| Suite | Upstream (T4, PyTorch) | Core ML (M5 Pro) |
|
| 78 |
+
| --- | ---: | ---: |
|
| 79 |
+
| jev.ag_news | 0.930 | **0.935** |
|
| 80 |
+
| jev.emotion | 0.530 | **0.537** |
|
| 81 |
+
| massive_intent.en | 0.657 | **0.657** |
|
| 82 |
+
| app.support_triage | 0.522 | **0.542** |
|
| 83 |
+
| app.email_spam | 0.993 | **0.993** |
|
| 84 |
+
| app.phishing | 0.993 | **0.993** |
|
| 85 |
+
| app.guardrails_jailbreak | 0.755 | **0.805** |
|
| 86 |
+
| app.moderation_toxicity | 0.525 | **0.535** |
|
| 87 |
+
| app.rag_relevance | 0.657 | **0.672** |
|
| 88 |
+
| app.model_routing_domain | 0.123 | **0.441** |
|
| 89 |
+
|
| 90 |
Conversion pipeline, verification reports, and Swift parity fixtures:
|
| 91 |
[mobius `models/computer-use/laya/coreml`](https://github.com/FluidInference/mobius).
|
| 92 |
|