Qwen3.6-35B-A3B: MLX 4-bit with native MTP
A 4-bit MLX checkpoint of Qwen/Qwen3.6-35B-A3B that keeps the model's native multi-token prediction (MTP) layer, so an engine that drafts with it gets everything from one repo.
What's in it
- The four
model-*.safetensorsshards and every config and tokenizer file are byte-identical to mlx-community/Qwen3.6-35B-A3B-4bit at revision38740b8: affine 4-bit, group size 64, routers at 8 bits. mtp-4bit.safetensorsadds the MTP layer that MLX conversions drop. It holds the officialmtp.*weights from Qwen/Qwen3.6-35B-A3B at revision995ad96, quantized the same way: affine 4-bit, group size 64, the router and shared-expert gate at 8 bits, norms in BF16. The tensors are namedlanguage_model.mtp.*.model.safetensors.index.jsondoesn't list the MTP file, so loaders that don't draft ignore it and load exactly the mlx-community model.
Use
With mlx-vlm, the same as the mlx-community conversion:
pip install -U mlx-vlm
python -m mlx_vlm.generate --model Vontra/Qwen3.6-35B-A3B-MLX-4bit-MTP --max-tokens 100 --temperature 0.0 --prompt "Describe this image." --image <path_to_image>
TensorFold drafts with the MTP layer and verifies every drafted token against the model, so its output equals the model's own decoding. Support for this model there is in development.
License
Apache-2.0, from Qwen/Qwen3.6-35B-A3B. The model is by the Qwen team; the MLX conversion of the main weights is by mlx-community.