LibertAIDAI/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF

🤗 Hugging Face 来源image-text-to-textapache-2.0激活 3B68 GBGGUF✓ 5 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo LibertAIDAI/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF ./model-folder
需要做种者 →

Qwen3.6-35B-A3B NVFP4 GGUF — MTP variant

NVFP4 GGUF quantizations of Qwen/Qwen3.6-35B-A3B with Multi-Token Prediction (MTP) support for speculative decoding in llama.cpp.

Same NVFP4 expert layout as our Qwen3.6-35B-A3B-NVFP4-GGUF repo, plus the MTP draft head extracted from the source for use with --spec-type draft-mtp.

Honest note up front: on the RTX 5090 with current llama.cpp, MTP does not yet win on this MoE for single-stream generation — the base MoE is already so fast (~250 tok/s with NVFP4-Q4_K_M) that draft + verify overhead exceeds the savings. The tuned config below gets it to break-even (-2%); the default loses ~27%. We're publishing the files anyway because (a) the MTP path is rapidly being optimized upstream and will likely flip, and (b) MTP weights are useful for downstream training/research. See the Performance section for numbers.

About LibertAI

LibertAI is a decentralized AI platform — private inference, an OpenAI-compatible API, and a chat UI, all running on community GPUs over Aleph Cloud instead of a single company's servers.

If you want to put this model (or any other) to work as an autonomous agent without running your own infrastructure, check out LiberClaw — Hermes-style agents hosted on Aleph Cloud with LibertAI inference.

Files

File Size Experts Other tensors When to pick
Qwen3.6-35B-A3B-NVFP4-Q4_K_M-mtp.gguf 19 GB NVFP4 Q4_K_M Recommended trunk
Qwen3.6-35B-A3B-NVFP4-Q8_0-mtp.gguf 20 GB NVFP4 Q8_0 Higher precision attention/embeddings
Qwen3.6-35B-A3B-NVFP4-BF16-mtp.gguf 22 GB NVFP4 BF16 Max source-fidelity for non-expert tensors
mtp-Qwen3.6-35B-A3B-NVFP4.gguf 3.5 GB BF16 (MoE) BF16/F32 MTP draft head — required for --spec-type draft-mtp. Itself a small MoE block (256 experts, 1 layer) plus shared experts and the nextn heads
mmproj-Qwen3.6-35B-A3B-F16.gguf 889 MB — F16 vision tower Required for image/video input

The trunks are built with convert_hf_to_gguf.py --no-mtp. The MTP weights are split into the separate mtp-*.gguf file so they can be loaded as a draft via --model-draft.

Performance

Measured on an NVIDIA RTX 5090 (32 GB, Blackwell, sm_120, CUDA 13.0), llama.cpp build dbe9c0c8c.

Single-stream serving, 28-token prompt → 512-token completion (n=3, prod settings: -c 262144 -fa on -ctk q4_0 -ctv q4_0 -ub 256 -b 1024 --parallel 1):

Config TG (tok/s) Draft accept rate Δ vs no-MTP
NVFP4-Q4_K_M, no MTP 252.9 — baseline (fastest)
NVFP4-Q4_K_M + MTP (n-max=2, p-min=0.3) 248.6 63.5% −2%
NVFP4-Q4_K_M + MTP (n-max=4, p-min=0.5, default) 184.8 65.6% −27%

The MoE base is so fast that MTP currently doesn't pay for itself at single-stream. With the tuned aggressive config (n-max=2 p-min=0.3), we get within 2% of baseline. We expect this to flip once the upstream llama.cpp MoE+MTP code path is optimized — the same was true for plain NVFP4 on MoE before recent improvements landed.

Until then: if you don't have a strong reason to use MTP on this MoE, prefer the non-MTP repo.

Usage

Server with MTP (experimental; tuned config)

llama-server \
  -m Qwen3.6-35B-A3B-NVFP4-Q4_K_M-mtp.gguf \
  --model-draft mtp-Qwen3.6-35B-A3B-NVFP4.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 2 \
  --spec-draft-p-min 0.3 \
  -ngl 999 -ngld 999 \
  -fa on -c 32768 \
  --host 0.0.0.0 --port 8080

Server without MTP (currently fastest single-stream)

llama-server \
  -m Qwen3.6-35B-A3B-NVFP4-Q4_K_M-mtp.gguf \
  -ngl 999 \
  -fa on -c 32768 \
  --host 0.0.0.0 --port 8080

Multimodal

Add --mmproj mmproj-Qwen3.6-35B-A3B-F16.gguf to either of the above for image/video input.

Recommended sampler

Qwen3.6 is a thinking model. For non-thinking usage, set chat_template_kwargs.enable_thinking=false in the API.

What is MTP?

Multi-Token Prediction adds a small extra block (here, 1 MoE layer at index 40) that produces a candidate next-token distribution from the same hidden state as the trunk. llama.cpp uses this as a draft for speculative decoding: generate up to N candidates, verify them in a single trunk forward pass, keep the accepted ones.

The MTP path landed in llama.cpp via PR #22673; NVFP4 scale-tensor support for the MTP block in MoE models landed in #23563.

Architecture notes

Qwen3.6-35B-A3B is an MoE model with 256 routed experts (8 active per token) plus a shared expert, hybrid attention + SSM (full attention every 4th of 40 layers). The NVFP4 source from mmangkad quantizes the routed expert FFN (ffn_{down,gate,up}_exps, 120 tensors = 40 layers × 3) — attention, SSM, and shared experts stay at higher precision. The MTP block at layer 40 is itself a small MoE layer with the same 256-expert structure (kept BF16 since the source didn't quantize it).

Sources & credits

License

Apache 2.0, inherited from the upstream model.