ddalcu/Qwen3.6-35B-A3B-MLX-Serve-4bit

🤗 Hugging Face 来源text-generationapache-2.035.1B 参数激活 3B70 GBsafetensors✓ 6 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo ddalcu/Qwen3.6-35B-A3B-MLX-Serve-4bit ./model-folder
需要做种者 →

Qwen3.6-35B-A3B for mlx-serve (4-bit, with MTP head)

The trunk of mlx-community/Qwen3.6-35B-A3B-4bit, unchanged, plus the MTP head from mlx-community/Qwen3.6-35B-A3B-MTP-4bit as the mlx-serve sidecar mtp/weights.safetensors. The head's tensors were renamed from the bare fc.weight layout to mtp.fc.weight; that is the only change, the bytes are identical.

mlx-serve --model ddalcu/Qwen3.6-35B-A3B-MLX-Serve-4bit --serve --mtp

MTP is opt-in on MoE trunks, hence --mtp. Without it the pack serves as a plain 4-bit Qwen3.6-35B-A3B.

Speed

M4 Max 128 GB, mlx-serve 26.9.2, tests/bench.sh:

serial MTP
decode (llmprobe bench) 166 tok/s 244 tok/s
predictable (code) 166 334
novel (prose) 165 168
16k context 122 177
4 concurrent, aggregate 152 134

Speculative decoding pays where the next tokens are predictable. On prose it is a wash, and under concurrency it costs you, because a batched decode already uses the width the draft rounds would have taken. Numbers from one box; yours will differ.

License

Apache-2.0, inherited from Qwen3.6-35B-A3B.