ddalcu/Qwen-Image-2.1-MLX-Serve-8bit

🤗 Hugging Face sourcetext-to-imageapache-2.021 GBsafetensors✓ 8 checksumsupdated today
Magnet🌱 0

Qwen-Image-2.1 MLX-Serve 8-bit

8-bit pack of Qwen/Qwen-Image-2.1 for mlx-serve: 18.4 GB, for 32 GB Macs.

What is in it

The checkpoint's own diffusers layout and key names, with the DiT block linears and the text-encoder layer linears affine-quantized to 8-bit (group 64). Kept dense: the VAE (f32), embed_tokens, norms, and the DiT's small or shared linears. Kept: the Qwen3-VL vision tower, for instruction editing. Dropped: lm_head and the VAE's per-frame time_convs. Built by tests/convert_qwen_image21_weights.py --preset 32gb.

Measured (M1 Pro, 32 GB)

Pack Size Steps Wall clock incl. load Peak memory
8-bit 1024x1024 40 985 s (~23 s/step) 12.95 GB
4-bit 1024x1024 3 87 s 9.55 GB
4-bit 512x512 20 118 s -

On a Mac the full set would crowd, mlx-serve loads the text encoder per request and frees it before the denoise, so the resident set is the DiT and VAE.

Run it

brew tap ddalcu/mlx-serve https://github.com/ddalcu/mlx-serve
brew install mlx-serve
mlx-serve pull ddalcu/Qwen-Image-2.1-MLX-Serve-8bit
mlx-serve serve
curl localhost:11234/v1/images/generations -H 'Content-Type: application/json' \
  -d '{"model":"ddalcu/Qwen-Image-2.1-MLX-Serve-8bit","prompt":"a red fox in fresh snow","size":"1024x1024"}'

40 steps when steps is omitted. guidance_scale above 1 with a negative_prompt runs real CFG (two forwards per step). image + strength does image-to-image. "mode":"edit" with an image (plus up to 9 ref_images) edits it from the prompt.

Apache-2.0, same as the base model.