Qwen3.6-35B-A3B NVFP4 GGUF — MTP variant
NVFP4 GGUF quantizations of Qwen/Qwen3.6-35B-A3B with Multi-Token Prediction (MTP) support for speculative decoding in llama.cpp.
Same NVFP4 expert layout as our Qwen3.6-35B-A3B-NVFP4-GGUF repo, plus the MTP draft head extracted from the source for use with --spec-type draft-mtp.
Honest note up front: on the RTX 5090 with current llama.cpp, MTP does not yet win on this MoE for single-stream generation — the base MoE is already so fast (~250 tok/s with NVFP4-Q4_K_M) that draft + verify overhead exceeds the savings. The tuned config below gets it to break-even (-2%); the default loses ~27%. We're publishing the files anyway because (a) the MTP path is rapidly being optimized upstream and will likely flip, and (b) MTP weights are useful for downstream training/research. See the Performance section for numbers.
About LibertAI
LibertAI is a decentralized AI platform — private inference, an OpenAI-compatible API, and a chat UI, all running on community GPUs over Aleph Cloud instead of a single company's servers.
If you want to put this model (or any other) to work as an autonomous agent without running your own infrastructure, check out LiberClaw — Hermes-style agents hosted on Aleph Cloud with LibertAI inference.
Files
| File | Size | Experts | Other tensors | When to pick |
|---|---|---|---|---|
Qwen3.6-35B-A3B-NVFP4-Q4_K_M-mtp.gguf |
19 GB | NVFP4 | Q4_K_M | Recommended trunk |
Qwen3.6-35B-A3B-NVFP4-Q8_0-mtp.gguf |
20 GB | NVFP4 | Q8_0 | Higher precision attention/embeddings |
Qwen3.6-35B-A3B-NVFP4-BF16-mtp.gguf |
22 GB | NVFP4 | BF16 | Max source-fidelity for non-expert tensors |
mtp-Qwen3.6-35B-A3B-NVFP4.gguf |
3.5 GB | BF16 (MoE) | BF16/F32 | MTP draft head — required for --spec-type draft-mtp. Itself a small MoE block (256 experts, 1 layer) plus shared experts and the nextn heads |
mmproj-Qwen3.6-35B-A3B-F16.gguf |
889 MB | — | F16 vision tower | Required for image/video input |
The trunks are built with convert_hf_to_gguf.py --no-mtp. The MTP weights are split into the separate mtp-*.gguf file so they can be loaded as a draft via --model-draft.
Performance
Measured on an NVIDIA RTX 5090 (32 GB, Blackwell, sm_120, CUDA 13.0), llama.cpp build dbe9c0c8c.
Single-stream serving, 28-token prompt → 512-token completion (n=3, prod settings: -c 262144 -fa on -ctk q4_0 -ctv q4_0 -ub 256 -b 1024 --parallel 1):
| Config | TG (tok/s) | Draft accept rate | Δ vs no-MTP |
|---|---|---|---|
| NVFP4-Q4_K_M, no MTP | 252.9 | — | baseline (fastest) |
| NVFP4-Q4_K_M + MTP (n-max=2, p-min=0.3) | 248.6 | 63.5% | −2% |
| NVFP4-Q4_K_M + MTP (n-max=4, p-min=0.5, default) | 184.8 | 65.6% | −27% |
The MoE base is so fast that MTP currently doesn't pay for itself at single-stream. With the tuned aggressive config (n-max=2 p-min=0.3), we get within 2% of baseline. We expect this to flip once the upstream llama.cpp MoE+MTP code path is optimized — the same was true for plain NVFP4 on MoE before recent improvements landed.
Until then: if you don't have a strong reason to use MTP on this MoE, prefer the non-MTP repo.
Usage
Server with MTP (experimental; tuned config)
llama-server \
-m Qwen3.6-35B-A3B-NVFP4-Q4_K_M-mtp.gguf \
--model-draft mtp-Qwen3.6-35B-A3B-NVFP4.gguf \
--spec-type draft-mtp \
--spec-draft-n-max 2 \
--spec-draft-p-min 0.3 \
-ngl 999 -ngld 999 \
-fa on -c 32768 \
--host 0.0.0.0 --port 8080
Server without MTP (currently fastest single-stream)
llama-server \
-m Qwen3.6-35B-A3B-NVFP4-Q4_K_M-mtp.gguf \
-ngl 999 \
-fa on -c 32768 \
--host 0.0.0.0 --port 8080
Multimodal
Add --mmproj mmproj-Qwen3.6-35B-A3B-F16.gguf to either of the above for image/video input.
Recommended sampler
Qwen3.6 is a thinking model. For non-thinking usage, set chat_template_kwargs.enable_thinking=false in the API.
What is MTP?
Multi-Token Prediction adds a small extra block (here, 1 MoE layer at index 40) that produces a candidate next-token distribution from the same hidden state as the trunk. llama.cpp uses this as a draft for speculative decoding: generate up to N candidates, verify them in a single trunk forward pass, keep the accepted ones.
The MTP path landed in llama.cpp via PR #22673; NVFP4 scale-tensor support for the MTP block in MoE models landed in #23563.
Architecture notes
Qwen3.6-35B-A3B is an MoE model with 256 routed experts (8 active per token) plus a shared expert, hybrid attention + SSM (full attention every 4th of 40 layers). The NVFP4 source from mmangkad quantizes the routed expert FFN (ffn_{down,gate,up}_exps, 120 tensors = 40 layers × 3) — attention, SSM, and shared experts stay at higher precision. The MTP block at layer 40 is itself a small MoE layer with the same 256-expert structure (kept BF16 since the source didn't quantize it).
Sources & credits
- Base model: Qwen/Qwen3.6-35B-A3B by Alibaba Qwen team — Apache 2.0
- NVFP4 calibration source: mmangkad/Qwen3.6-35B-A3B-NVFP4 (NVIDIA ModelOpt v0.43)
- mmproj source: official BF16 weights from
Qwen/Qwen3.6-35B-A3B - Tooling: llama.cpp
convert_hf_to_gguf.py --no-mtp/--mtpandllama-quantize
License
Apache 2.0, inherited from the upstream model.