alexwengg commited on
Commit
012b42d
·
verified ·
1 Parent(s): 791c64e

Model card: benchmark table

Browse files
Files changed (1) hide show
  1. README.md +17 -0
README.md CHANGED
@@ -70,6 +70,23 @@ Apple M5 Pro, macOS 27.0, 16 fixture questions vs. the unmodified PyTorch FP32 r
70
  | L512 | 27.5 ms | **9.0 ms** |
71
  | L1024 | 80.1 ms | **17.9 ms** |
72
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
73
  Conversion pipeline, verification reports, and Swift parity fixtures:
74
  [mobius `models/computer-use/laya/coreml`](https://github.com/FluidInference/mobius).
75
 
 
70
  | L512 | 27.5 ms | **9.0 ms** |
71
  | L1024 | 80.1 ms | **17.9 ms** |
72
 
73
+ On laya's published application suites (3,899 questions, seed 13, rebuilt from upstream's scripts),
74
+ the Core ML buckets answered from Swift match the PyTorch reference's accuracy on every suite at
75
+ 5.2 ms median per question (p95 18 ms):
76
+
77
+ | Suite | Upstream (T4, PyTorch) | Core ML (M5 Pro) |
78
+ | --- | ---: | ---: |
79
+ | jev.ag_news | 0.930 | **0.935** |
80
+ | jev.emotion | 0.530 | **0.537** |
81
+ | massive_intent.en | 0.657 | **0.657** |
82
+ | app.support_triage | 0.522 | **0.542** |
83
+ | app.email_spam | 0.993 | **0.993** |
84
+ | app.phishing | 0.993 | **0.993** |
85
+ | app.guardrails_jailbreak | 0.755 | **0.805** |
86
+ | app.moderation_toxicity | 0.525 | **0.535** |
87
+ | app.rag_relevance | 0.657 | **0.672** |
88
+ | app.model_routing_domain | 0.123 | **0.441** |
89
+
90
  Conversion pipeline, verification reports, and Swift parity fixtures:
91
  [mobius `models/computer-use/laya/coreml`](https://github.com/FluidInference/mobius).
92