File size: 2,693 Bytes
f3cedaa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
# MLX is not shipped for this model

The MLX build was measured and rejected. The exporter in `export-scripts` is
kept so the decision can be re-tested against a future ExecuTorch/MLX release,
but no `.pte` or `config.json` is published here.

## Why: it is slower, and its output is not finite

Measured on device (iPhone 16, ExecuTorch 1.4.1, 3 warmup runs + 15 timed runs,
median, inputs allocated from each model's own schema and filled with 0.5). All
four styles share one architecture and differ only in weights, so all four were
run rather than measuring one and generalising:

| style | Core ML fp16 | MLX fp16 | MLX penalty | MLX output |
| --- | --- | --- | --- | --- |
| candy | **193.4 ms** | 564.5 ms | 2.92x slower | non-finite |
| mosaic | **194.7 ms** | 546.4 ms | 2.81x slower | non-finite |
| rain princess | **193.3 ms** | 566.7 ms | 2.93x slower | non-finite |
| udnie | **196.9 ms** | 574.7 ms | 2.92x slower | non-finite |

The latency alone would be reason enough, but the output settles it. Every MLX
build emitted NaN or Inf, while the Core ML build fed the identical input on the
same device in the same run returned finite output in all four cases. This
artifact does not merely run slowly, it produces garbage.

Size is not an argument either. MLX fp16 is 3.48 MB against Core ML fp16 at
3.79 MB, a 0.3 MB difference on a model of this size.

## Root cause: op coverage

The ExecuTorch MLX backend has no handler for three ops this architecture uses
everywhere, so the graph shatters into **50 subgraphs**:

| unsupported op | count | where it comes from |
| --- | --- | --- |
| `aten.flip` | 64 | the reflection pads |
| `aten._native_batch_norm_legit.no_stats` | 15 | InstanceNorm |
| `aten.upsample_bilinear2d` | 2 | the two upsampling stages |

Note the specific batch-norm overload. The MLX backend does register
`_native_batch_norm_legit_no_training`, but not the `no_stats` variant that
InstanceNorm lowers to, and a style transfer network is InstanceNorm in every
block.

This is the same class of failure as the SDXS text-to-image model, which
fragments into 28 subgraphs for want of `native_group_norm`, filed upstream as
https://github.com/pytorch/executorch/issues/22017.

## Why style transfer is not a special case

MLX runs on the Metal GPU; Core ML lowers to the ANE, which is purpose-built for
this kind of network. MLX was measured behind Core ML on every convolutional
vision model tested. Its real advantage is transformer workloads, where
weight-only quantization shrinks the artifact substantially: CLIP vision is 97%
`nn.Linear` and goes 351 MB to 98 MB. A style transfer network is pure
convolution and lands on the wrong side of that line.