LibertAIDAI/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF

🤗 Hugging Face sourceimage-text-to-textapache-2.03B activated68 GBGGUF✓ 5 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo LibertAIDAI/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF ./model-folder
Needs a seeder →

Qwen3.6-35B-A3B NVFP4 GGUF — MTP variant

NVFP4 GGUF quantizations of Qwen/Qwen3.6-35B-A3B with Multi-Token Prediction (MTP) support for speculative decoding in llama.cpp.

Same NVFP4 expert layout as our Qwen3.6-35B-A3B-NVFP4-GGUF repo, plus the MTP draft head extracted from the source for use with --spec-type draft-mtp.

Honest note up front: on the RTX 5090 with current llama.cpp, MTP does not yet win on this MoE for single-stream generation — the base MoE is already so fast (~250 tok/s with NVFP4-Q4_K_M) that draft + verify overhead exceeds the savings. The tuned config below gets it to break-even (-2%); the default loses ~27%. We're publishing the files anyway because (a) the MTP path is rapidly being optimized upstream and will likely flip, and (b) MTP weights are useful for downstream training/research. See the Performance section for numbers.

About LibertAI

LibertAI is a decentralized AI platform — private inference, an OpenAI-compatible API, and a chat UI, all running on community GPUs over Aleph Cloud instead of a single company's servers.

If you want to put this model (or any other) to work as an autonomous agent without running your own infrastructure, check out LiberClaw — Hermes-style agents hosted on Aleph Cloud with LibertAI inference.

Files

File Size Experts Other tensors When to pick
Qwen3.6-35B-A3B-NVFP4-Q4_K_M-mtp.gguf 19 GB NVFP4 Q4_K_M Recommended trunk
Qwen3.6-35B-A3B-NVFP4-Q8_0-mtp.gguf 20 GB NVFP4 Q8_0 Higher precision attention/embeddings
Qwen3.6-35B-A3B-NVFP4-BF16-mtp.gguf 22 GB NVFP4 BF16 Max source-fidelity for non-expert tensors
mtp-Qwen3.6-35B-A3B-NVFP4.gguf 3.5 GB BF16 (MoE) BF16/F32 MTP draft head — required for --spec-type draft-mtp. Itself a small MoE block (256 experts, 1 layer) plus shared experts and the nextn heads
mmproj-Qwen3.6-35B-A3B-F16.gguf 889 MB — F16 vision tower Required for image/video input

The trunks are built with convert_hf_to_gguf.py --no-mtp. The MTP weights are split into the separate mtp-*.gguf file so they can be loaded as a draft via --model-draft.

Performance

Measured on an NVIDIA RTX 5090 (32 GB, Blackwell, sm_120, CUDA 13.0), llama.cpp build dbe9c0c8c.

Single-stream serving, 28-token prompt → 512-token completion (n=3, prod settings: -c 262144 -fa on -ctk q4_0 -ctv q4_0 -ub 256 -b 1024 --parallel 1):

Config TG (tok/s) Draft accept rate Δ vs no-MTP
NVFP4-Q4_K_M, no MTP 252.9 — baseline (fastest)
NVFP4-Q4_K_M + MTP (n-max=2, p-min=0.3) 248.6 63.5% −2%
NVFP4-Q4_K_M + MTP (n-max=4, p-min=0.5, default) 184.8 65.6% −27%

The MoE base is so fast that MTP currently doesn't pay for itself at single-stream. With the tuned aggressive config (n-max=2 p-min=0.3), we get within 2% of baseline. We expect this to flip once the upstream llama.cpp MoE+MTP code path is optimized — the same was true for plain NVFP4 on MoE before recent improvements landed.

Until then: if you don't have a strong reason to use MTP on this MoE, prefer the non-MTP repo.

Usage

Server with MTP (experimental; tuned config)

llama-server \
  -m Qwen3.6-35B-A3B-NVFP4-Q4_K_M-mtp.gguf \
  --model-draft mtp-Qwen3.6-35B-A3B-NVFP4.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 2 \
  --spec-draft-p-min 0.3 \
  -ngl 999 -ngld 999 \
  -fa on -c 32768 \
  --host 0.0.0.0 --port 8080

Server without MTP (currently fastest single-stream)

llama-server \
  -m Qwen3.6-35B-A3B-NVFP4-Q4_K_M-mtp.gguf \
  -ngl 999 \
  -fa on -c 32768 \
  --host 0.0.0.0 --port 8080

Multimodal

Add --mmproj mmproj-Qwen3.6-35B-A3B-F16.gguf to either of the above for image/video input.

Recommended sampler

Qwen3.6 is a thinking model. For non-thinking usage, set chat_template_kwargs.enable_thinking=false in the API.

What is MTP?

Multi-Token Prediction adds a small extra block (here, 1 MoE layer at index 40) that produces a candidate next-token distribution from the same hidden state as the trunk. llama.cpp uses this as a draft for speculative decoding: generate up to N candidates, verify them in a single trunk forward pass, keep the accepted ones.

The MTP path landed in llama.cpp via PR #22673; NVFP4 scale-tensor support for the MTP block in MoE models landed in #23563.

Architecture notes

Qwen3.6-35B-A3B is an MoE model with 256 routed experts (8 active per token) plus a shared expert, hybrid attention + SSM (full attention every 4th of 40 layers). The NVFP4 source from mmangkad quantizes the routed expert FFN (ffn_{down,gate,up}_exps, 120 tensors = 40 layers × 3) — attention, SSM, and shared experts stay at higher precision. The MTP block at layer 40 is itself a small MoE layer with the same 256-expert structure (kept BF16 since the source didn't quantize it).

Sources & credits

License

Apache 2.0, inherited from the upstream model.