Fine-tuned Qwen 2.5 (0.5B β 3B) on real coding-agent traces, 10 controlled runs, one 16GB Mac. Compared PyTorch MPS vs. Apple MLX for local LoRA SFT β and the honest answer is "it depends on what you're optimizing for":
β’ PyTorch MPS: 2.2xβ5.7x faster raw throughput, but hits a hard memory wall β can't load a 3B model in FP16 on 16GB. β’ Apple MLX: 4-bit QLoRA fits 3B+ models with almost flat memory scaling as context grows (+109 MB going from 1kβ4k tokens). β’ 4-bit quantization doesn't cost you convergence β eval loss tracks closely across backends. β’ The bigger surprise: most of MLX's slowdown isn't the 4-bit dequant tax. Two of the 10 runs went unquantized to isolate it β dequant only explains 1.07xβ1.4x of the gap. A ~4.1β4.6x framework-level gap remains either way.
All 10 LoRA adapters + Trackio logs are public so the numbers are checkable, not just claimed.
Saw it being used as a computer automation tool (i.e. open an app, type, something, etc) I think once more open jev type models come out we will have crazy things