Qwen-Image-2.1 MLX-Serve 8-bit
8-bit pack of Qwen/Qwen-Image-2.1 for mlx-serve: 18.4 GB, for 32 GB Macs.
What is in it
The checkpoint's own diffusers layout and key names, with the DiT block linears and the
text-encoder layer linears affine-quantized to 8-bit (group 64). Kept dense: the VAE
(f32), embed_tokens, norms, and the DiT's small or shared linears. Kept: the Qwen3-VL
vision tower, for instruction editing. Dropped: lm_head and the VAE's per-frame time_convs.
Built by tests/convert_qwen_image21_weights.py --preset 32gb.
Measured (M1 Pro, 32 GB)
| Pack | Size | Steps | Wall clock incl. load | Peak memory |
|---|---|---|---|---|
| 8-bit | 1024x1024 | 40 | 985 s (~23 s/step) | 12.95 GB |
| 4-bit | 1024x1024 | 3 | 87 s | 9.55 GB |
| 4-bit | 512x512 | 20 | 118 s | - |
On a Mac the full set would crowd, mlx-serve loads the text encoder per request and frees it before the denoise, so the resident set is the DiT and VAE.
Run it
brew tap ddalcu/mlx-serve https://github.com/ddalcu/mlx-serve
brew install mlx-serve
mlx-serve pull ddalcu/Qwen-Image-2.1-MLX-Serve-8bit
mlx-serve serve
curl localhost:11234/v1/images/generations -H 'Content-Type: application/json' \
-d '{"model":"ddalcu/Qwen-Image-2.1-MLX-Serve-8bit","prompt":"a red fox in fresh snow","size":"1024x1024"}'
40 steps when steps is omitted. guidance_scale above 1 with a negative_prompt runs
real CFG (two forwards per step). image + strength does image-to-image. "mode":"edit" with an image (plus up to 9 ref_images)
edits it from the prompt.
Apache-2.0, same as the base model.