Qwen3.6-35B-A3B for mlx-serve (4-bit, with MTP head)
The trunk of
mlx-community/Qwen3.6-35B-A3B-4bit,
unchanged, plus the MTP head from
mlx-community/Qwen3.6-35B-A3B-MTP-4bit
as the mlx-serve sidecar mtp/weights.safetensors. The head's tensors were
renamed from the bare fc.weight layout to mtp.fc.weight; that is the only
change, the bytes are identical.
mlx-serve --model ddalcu/Qwen3.6-35B-A3B-MLX-Serve-4bit --serve --mtp
MTP is opt-in on MoE trunks, hence --mtp. Without it the pack serves as a
plain 4-bit Qwen3.6-35B-A3B.
Speed
M4 Max 128 GB, mlx-serve 26.9.2, tests/bench.sh:
| serial | MTP | |
|---|---|---|
| decode (llmprobe bench) | 166 tok/s | 244 tok/s |
| predictable (code) | 166 | 334 |
| novel (prose) | 165 | 168 |
| 16k context | 122 | 177 |
| 4 concurrent, aggregate | 152 | 134 |
Speculative decoding pays where the next tokens are predictable. On prose it is a wash, and under concurrency it costs you, because a batched decode already uses the width the draft rounds would have taken. Numbers from one box; yours will differ.
License
Apache-2.0, inherited from Qwen3.6-35B-A3B.