ngquocvinh/Qwen3.8-35B-A3B-Distill-GGUF

🤗 Hugging Face 来源text-generationapache-2.0激活 3B156 GBGGUF✓ 9 个校验和今天更新
需要做种者 →

Qwen3.8-35B-A3B-Distill GGUF

Community GGUF quantizations of empero-ai/Qwen3.8-35B-A3B-Distill.

☕ If this GGUF made your day easier, a coffee would make mine.
Send a coffee ☕
I build and test these releases myself. Your coffee helps keep me going.
Thank you for supporting this work.

About Qwen3.8-35B-A3B-Distill

This release is based on the upstream Qwen3.8-35B-A3B-Distill model card. The model is a text-generation MoE model with 35B total parameters and approximately 3B active parameters per token, using 256 experts with 8 routed experts. It uses hybrid attention components and has a native context length of 262,144 tokens in the upstream configuration.

The GGUF files are intended for local inference with llama.cpp and compatible runtimes. The upstream model card and runtime support should be consulted before using long-context, tool-calling, or other task-specific features.

No training or fine-tuning was performed for this release. These are community GGUF quantizations, not an official Qwen3.8 release or endorsement by the upstream authors.

Files

File Approx. size Notes
Qwen3.8-35B-A3B-Distill-Q8_0.gguf 35.22 GiB Highest published precision in this package
Qwen3.8-35B-A3B-Distill-Q6_K.gguf 27.20 GiB High-fidelity option
Qwen3.8-35B-A3B-Distill-Q5_K_M.gguf 23.60 GiB High-fidelity / memory-balanced option
Qwen3.8-35B-A3B-Distill-Q4_K_M.gguf 20.22 GiB Balanced default candidate
Qwen3.8-35B-A3B-Distill-Q3_K_M.gguf 15.98 GiB Experimental lower-bit option
Qwen3.8-35B-A3B-Distill-Q2_K.gguf 12.34 GiB Experimental lower-bit option
Qwen3.8-35B-A3B-Distill-IQ2_XS.gguf 10.01 GiB Experimental ultra-low-bit option

Fidelity measurements

The measurements below compare each published GGUF with the locked BF16 GGUF source. They are next-token fidelity measurements, not a general capability benchmark. Values are averaged over WikiText-2 valid and test, using 2 chunks per split, context length 2048, and the same llama.cpp runtime. Lower KLD, ΔPPL, and RMS Δp indicate closer next-token behavior to BF16; higher Top-1 agreement is better.

File Mean KLD ↓ Top-1 vs BF16 ↑ ΔPPL RMS Δp
Qwen3.8-35B-A3B-Distill-Q8_0.gguf 0.002454 97.972% -0.068% 1.396%
Qwen3.8-35B-A3B-Distill-Q6_K.gguf 0.004134 97.166% -0.226% 1.793%
Qwen3.8-35B-A3B-Distill-Q5_K_M.gguf 0.008311 95.748% +0.205% 2.660%
Qwen3.8-35B-A3B-Distill-Q4_K_M.gguf 0.017029 94.135% +1.299% 4.032%
Qwen3.8-35B-A3B-Distill-Q3_K_M.gguf 0.052876 90.274% +4.728% 6.933%
Qwen3.8-35B-A3B-Distill-Q2_K.gguf 0.125190 84.213% +12.617% 10.822%
Qwen3.8-35B-A3B-Distill-IQ2_XS.gguf 0.234263 79.326% +28.715% 15.648%

For a general memory/quality balance, Q4_K_M is the practical starting point from this measurement set. Q3_K_M, Q2_K, and IQ2_XS should be treated as experimental choices where memory limits justify the measured fidelity trade-off. IQ2_XS is the smallest published option in this update and shows a substantially larger fidelity deviation from BF16. Quality summaries and exact reproduction details are in reproducibility/quality-summary.tsv and reproducibility/manifest.md.

Quick start

From a llama.cpp checkout, with reproducibility/chat_template.jinja downloaded from this repository:

./llama-cli \
  -m Qwen3.8-35B-A3B-Distill-Q4_K_M.gguf \
  --chat-template-file reproducibility/chat_template.jinja \
  --jinja \
  --reasoning off \
  --cpu-moe \
  -p 'Answer briefly in English: What is GGUF and why is it useful for running language models locally?' \
  -n 128 -c 4096 -ngl 8

Adjust -ngl for the available GPU memory. The validation profile for this release used 8 GPU layers, CPU MoE, and llama.cpp's qwen35moe support on an NVIDIA A10M.

Reproducibility and validation

Raw conversion, quantization, smoke-test, fidelity, and benchmark logs are kept locally under reports/ and are not part of this public package.

License and attribution

The upstream model is released under the Apache-2.0 license. See LICENSE and the upstream model card for the original model's terms and attribution.

These are community GGUF quantizations, not an official Qwen3.8 release or endorsement.