msluszniak's picture
Document why MLX is not shipped: 2.9x slower than Core ML and non-finite output
f3cedaa verified
|
Raw
History Blame Contribute Delete
2.69 kB

MLX is not shipped for this model

The MLX build was measured and rejected. The exporter in export-scripts is kept so the decision can be re-tested against a future ExecuTorch/MLX release, but no .pte or config.json is published here.

Why: it is slower, and its output is not finite

Measured on device (iPhone 16, ExecuTorch 1.4.1, 3 warmup runs + 15 timed runs, median, inputs allocated from each model's own schema and filled with 0.5). All four styles share one architecture and differ only in weights, so all four were run rather than measuring one and generalising:

style Core ML fp16 MLX fp16 MLX penalty MLX output
candy 193.4 ms 564.5 ms 2.92x slower non-finite
mosaic 194.7 ms 546.4 ms 2.81x slower non-finite
rain princess 193.3 ms 566.7 ms 2.93x slower non-finite
udnie 196.9 ms 574.7 ms 2.92x slower non-finite

The latency alone would be reason enough, but the output settles it. Every MLX build emitted NaN or Inf, while the Core ML build fed the identical input on the same device in the same run returned finite output in all four cases. This artifact does not merely run slowly, it produces garbage.

Size is not an argument either. MLX fp16 is 3.48 MB against Core ML fp16 at 3.79 MB, a 0.3 MB difference on a model of this size.

Root cause: op coverage

The ExecuTorch MLX backend has no handler for three ops this architecture uses everywhere, so the graph shatters into 50 subgraphs:

unsupported op count where it comes from
aten.flip 64 the reflection pads
aten._native_batch_norm_legit.no_stats 15 InstanceNorm
aten.upsample_bilinear2d 2 the two upsampling stages

Note the specific batch-norm overload. The MLX backend does register _native_batch_norm_legit_no_training, but not the no_stats variant that InstanceNorm lowers to, and a style transfer network is InstanceNorm in every block.

This is the same class of failure as the SDXS text-to-image model, which fragments into 28 subgraphs for want of native_group_norm, filed upstream as https://github.com/pytorch/executorch/issues/22017.

Why style transfer is not a special case

MLX runs on the Metal GPU; Core ML lowers to the ANE, which is purpose-built for this kind of network. MLX was measured behind Core ML on every convolutional vision model tested. Its real advantage is transformer workloads, where weight-only quantization shrinks the artifact substantially: CLIP vision is 97% nn.Linear and goes 351 MB to 98 MB. A style transfer network is pure convolution and lands on the wrong side of that line.