Qwen3.8-35B-A3B-Distill GGUF
Community GGUF quantizations of empero-ai/Qwen3.8-35B-A3B-Distill.
☕ If this GGUF made your day easier, a coffee would make mine.Send a coffee ☕
I build and test these releases myself. Your coffee helps keep me going.
Thank you for supporting this work.
About Qwen3.8-35B-A3B-Distill
This release is based on the upstream Qwen3.8-35B-A3B-Distill model card. The model is a text-generation MoE model with 35B total parameters and approximately 3B active parameters per token, using 256 experts with 8 routed experts. It uses hybrid attention components and has a native context length of 262,144 tokens in the upstream configuration.
The GGUF files are intended for local inference with llama.cpp and compatible runtimes. The upstream model card and runtime support should be consulted before using long-context, tool-calling, or other task-specific features.
No training or fine-tuning was performed for this release. These are community GGUF quantizations, not an official Qwen3.8 release or endorsement by the upstream authors.
Files
| File | Approx. size | Notes |
|---|---|---|
Qwen3.8-35B-A3B-Distill-Q8_0.gguf |
35.22 GiB | Highest published precision in this package |
Qwen3.8-35B-A3B-Distill-Q6_K.gguf |
27.20 GiB | High-fidelity option |
Qwen3.8-35B-A3B-Distill-Q5_K_M.gguf |
23.60 GiB | High-fidelity / memory-balanced option |
Qwen3.8-35B-A3B-Distill-Q4_K_M.gguf |
20.22 GiB | Balanced default candidate |
Qwen3.8-35B-A3B-Distill-Q3_K_M.gguf |
15.98 GiB | Experimental lower-bit option |
Qwen3.8-35B-A3B-Distill-Q2_K.gguf |
12.34 GiB | Experimental lower-bit option |
Qwen3.8-35B-A3B-Distill-IQ2_XS.gguf |
10.01 GiB | Experimental ultra-low-bit option |
Fidelity measurements
The measurements below compare each published GGUF with the locked BF16 GGUF source. They are next-token fidelity measurements, not a general capability benchmark. Values are averaged over WikiText-2 valid and test, using 2 chunks per split, context length 2048, and the same llama.cpp runtime. Lower KLD, ΔPPL, and RMS Δp indicate closer next-token behavior to BF16; higher Top-1 agreement is better.
| File | Mean KLD ↓ | Top-1 vs BF16 ↑ | ΔPPL | RMS Δp |
|---|---|---|---|---|
Qwen3.8-35B-A3B-Distill-Q8_0.gguf |
0.002454 | 97.972% | -0.068% | 1.396% |
Qwen3.8-35B-A3B-Distill-Q6_K.gguf |
0.004134 | 97.166% | -0.226% | 1.793% |
Qwen3.8-35B-A3B-Distill-Q5_K_M.gguf |
0.008311 | 95.748% | +0.205% | 2.660% |
Qwen3.8-35B-A3B-Distill-Q4_K_M.gguf |
0.017029 | 94.135% | +1.299% | 4.032% |
Qwen3.8-35B-A3B-Distill-Q3_K_M.gguf |
0.052876 | 90.274% | +4.728% | 6.933% |
Qwen3.8-35B-A3B-Distill-Q2_K.gguf |
0.125190 | 84.213% | +12.617% | 10.822% |
Qwen3.8-35B-A3B-Distill-IQ2_XS.gguf |
0.234263 | 79.326% | +28.715% | 15.648% |
For a general memory/quality balance, Q4_K_M is the practical starting point from this measurement set. Q3_K_M, Q2_K, and IQ2_XS should be treated as experimental choices where memory limits justify the measured fidelity trade-off. IQ2_XS is the smallest published option in this update and shows a substantially larger fidelity deviation from BF16. Quality summaries and exact reproduction details are in reproducibility/quality-summary.tsv and reproducibility/manifest.md.
Quick start
From a llama.cpp checkout, with reproducibility/chat_template.jinja downloaded from this repository:
./llama-cli \
-m Qwen3.8-35B-A3B-Distill-Q4_K_M.gguf \
--chat-template-file reproducibility/chat_template.jinja \
--jinja \
--reasoning off \
--cpu-moe \
-p 'Answer briefly in English: What is GGUF and why is it useful for running language models locally?' \
-n 128 -c 4096 -ngl 8
Adjust -ngl for the available GPU memory. The validation profile for this release used 8 GPU layers, CPU MoE, and llama.cpp's qwen35moe support on an NVIDIA A10M.
Reproducibility and validation
SHA256SUMS.txtcontains checksums for the public files.reproducibility/manifest.mdrecords the locked BF16 source, calibration data, imatrix, runtime, commands, and validation profile.RELEASE_MANIFEST.jsonrecords the machine-readable source and artifact ledger.reproducibility/quality-summary.tsvcontains the machine-readable fidelity table.reproducibility/runtime-summary.tsvcontains the optional Q4_K_M llama-bench result.reproducibility/chat_template.jinjais the upstream chat template used by the smoke test.
Raw conversion, quantization, smoke-test, fidelity, and benchmark logs are kept locally under reports/ and are not part of this public package.
License and attribution
The upstream model is released under the Apache-2.0 license. See LICENSE and the upstream model card for the original model's terms and attribution.
These are community GGUF quantizations, not an official Qwen3.8 release or endorsement.