ddalcu/Qwen-Image-2.1-MLX-Serve-4bit

🤗 Hugging Face sourcetext-to-imageapache-2.013 GBsafetensors✓ 8 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo ddalcu/Qwen-Image-2.1-MLX-Serve-4bit ./model-folder
Needs a seeder →

Qwen-Image-2.1 MLX-Serve 4-bit

4-bit pack of Qwen/Qwen-Image-2.1 for mlx-serve: 11.2 GB, for 16 GB Macs.

What is in it

The checkpoint's own diffusers layout and key names, with the DiT block linears and the text-encoder layer linears affine-quantized to 4-bit (group 64). Kept dense: the VAE (f32), embed_tokens, norms, and the DiT's small or shared linears. Kept: the Qwen3-VL vision tower, for instruction editing. Dropped: lm_head and the VAE's per-frame time_convs. Built by tests/convert_qwen_image21_weights.py --preset 16gb.

Measured (M1 Pro, 32 GB)

Pack Size Steps Wall clock incl. load Peak memory
8-bit 1024x1024 40 985 s (~23 s/step) 12.95 GB
4-bit 1024x1024 3 87 s 9.55 GB
4-bit 512x512 20 118 s -

On a Mac the full set would crowd, mlx-serve loads the text encoder per request and frees it before the denoise, so the resident set is the DiT and VAE.

Run it

brew tap ddalcu/mlx-serve https://github.com/ddalcu/mlx-serve
brew install mlx-serve
mlx-serve pull ddalcu/Qwen-Image-2.1-MLX-Serve-4bit
mlx-serve serve
curl localhost:11234/v1/images/generations -H 'Content-Type: application/json' \
  -d '{"model":"ddalcu/Qwen-Image-2.1-MLX-Serve-4bit","prompt":"a red fox in fresh snow","size":"1024x1024"}'

40 steps when steps is omitted. guidance_scale above 1 with a negative_prompt runs real CFG (two forwards per step). image + strength does image-to-image. "mode":"edit" with an image (plus up to 9 ref_images) edits it from the prompt.

Apache-2.0, same as the base model.