mlx-community/Ming-Image-0.1-Design-Layer-4bit

Pre-quantized MLX tier of inclusionAI/Ming-Image-0.1-Design-Layer (MIT) for Apple Silicon, loaded by the Swift/MLX port ming-image-swift. 4-bit tier of the layer-decomposition model, for memory-bound Macs. It fits a 36 GB Mac, and the decompositions still match bf16.

A Mac never has to hold the bf16 weights to use this tier. Total size: 20.5 GB; the bf16 repo is 49.8 GB. It is part of the Ming-Image (MLX) collection, next to mlx-community/Ming-Image-0.1-Design-Layer-bf16.

What is quantized

Component This tier
mllm/: MoE MLLM (attention, dense and shared-expert MLPs, the 256 routed experts) 4-bit
connector/: Qwen2-1.5B connector 8-bit
transformer/: DiT attention and feed-forward 8-bit
Kept at full precision: the MoE routers, embeddings, norms, the Qwen2.5 ViT, the f32 projection heads (mlp/), the VAE, and the DiT's conditioning layers (adaLN, embedders, final layer) bf16 / f32

Weight-only affine quantization, group size 64. The DiT never goes below 8 bits: weight-only int4 on a DiT is real quality damage that buys no speed.

The layout is the bf16 repo's upstream tree, with MLX .scales / .biases stored beside each quantized weight and a quantization block in each quantized component's config.json. The Swift loader reads it as published. Loading this repo gives parameters bit-identical to quantizing the bf16 snapshot at load time (verified for every parameter).

Quality

Gate: five real signage stills at 512, plus one at 1024 (same noise as bf16) bf16 8-bit 4-bit
Layers recomposited vs the input 24.4–33.7 dB within 0.1 dB within 0.2 dB of bf16 on every job
Layer coverage — identical identical
Stray text outside the text layer (the hardest still) 26 chars 30 35

Full tables are in GATE-RESULTS §8 in the port's oracle (https://github.com/xocialize/ming-image-swift).

Memory and speed

Measured on an M5 Max as process phys_footprint, with MLX's buffer cache capped at 2 GB (MLXEngine's default).

  • Post-load resident: 7.6 GB. Peak process footprint: 20.0–20.6 GB for 4 to 12 layers at the 1024 bucket. At 12 layers the multi-frame denoise overtakes conditioning.
  • MLXEngine declares 7.9 GB resident plus 15.6 GB activation, which is admitted on a 36 GB Mac. Requests are capped at 12 layers, the measured envelope.
  • Speed (M5 Max, 12 steps, CFG 2.0): 134 s for 4 layers at the 512 bucket. bf16 takes 131 s.

Use (Swift / MLXEngine)

import MLXMingImage
import MLXToolKit

let package = MingImageLayerPackage(configuration: MingImageLayerConfiguration(quant: .int4))
try await package.load()
let response = try await package.run(LayerDecomposeRequest(image: design, spec: spec, resolution: 1024)) as! LayerDecomposeResponse

Code: https://github.com/xocialize/ming-image-swift

License

MIT, as the upstream weights. The upstream LICENSE is included.

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/Ming-Image-0.1-Design-Layer-4bit

Finetuned
(4)
this model

Collection including mlx-community/Ming-Image-0.1-Design-Layer-4bit