Qwen3-30B-A3B — Pollard
Pollard shrank this model: 61.06 GB (f16) → 7.54 GB — 88% smaller, 8.1× down.
The smallest rung here; larger, higher-fidelity rungs are listed below.
format this model's size f16 61.06 GB Q8_0 ~32.36 GB Q6_K ~25.03 GB Q4_K_M ~17.71 GB PollardMix (this repo's IQ1_KT) 7.54 GB
Pollard builds of Qwen/Qwen3-30B-A3B made with Pollard Weights — a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).
Standard GGUF — runs in stock llama.cpp / ik_llama.cpp, Ollama, LM Studio. Trellis (IQ*_KT) files need ik_llama.cpp; the K-quants run anywhere.
Model details
| Parameter count | ~30.5B |
| Architecture | qwen3_moe |
| Input support | text |
| imatrix | no |
| Perplexity measured | yes — table below |
Which file should I choose?
Every rung is the same weights, sized to a different RAM budget by the measured allocation. Pick the largest one that fits your machine with room for context:
- ~27 GB RAM / VRAM →
Q6_K(25.12 GB). near-lossless - ~18 GB RAM / VRAM →
IQ4_XS(16.48 GB). recommended default - ~15 GB RAM / VRAM →
IQ3_S(13.46 GB). smaller - ~10 GB RAM / VRAM →
IQ1_KT(7.54 GB). flagship — 1-bit mixed-precision MoE trellis
Available files (WikiText-2 raw, ctx 2048, 145 chunks; KLD vs Q6_K base)
| file | PPL | size | Mean KLD | notes |
|---|---|---|---|---|
Qwen3-30B-A3B-Pollard-IQ1_KT.gguf |
— | 7.54 GB | — | flagship — 1-bit mixed-precision MoE trellis |
Qwen3-30B-A3B-Pollard-IQ3_S.gguf |
— | 13.46 GB | — | smaller |
Qwen3-30B-A3B-Pollard-IQ4_XS.gguf |
— | 16.48 GB | — | recommended default |
Qwen3-30B-A3B-Pollard-Q6_K.gguf |
— | 25.12 GB | — | near-lossless |
The numbers (WikiText-2 raw, ctx 2048, 145 chunks; KLD vs Q6_K base)
| build | role | PPL | size | bpw | Mean KLD | Median KLD | top-1 |
|---|---|---|---|---|---|---|---|
| uniform IQ2_KT | 2-bit ceiling | 7.28 | 8.34 GB | 2.19 | 0.134 | 0.059 | 84.81% |
| PollardMix | this model | 8.57 | 7.02 GB | 1.84 | 0.310 | 0.140 | 77.81% |
| uniform IQ1_KT | 1-bit baseline | 9.01 | 6.57 GB | 1.73 | 0.360 | 0.174 | 75.47% |
PollardMix beats the uniform 1-bit trellis quant on every metric — PPL −4.9%, Mean KLD −14%, Median KLD −20%, top-1 +2.3 pts — at +6.9% size, under the 2-bit ceiling. Same clean sweep as the 7B/14B dense cards, now reproduced on a MoE — the automap policy generalizes (crush cold experts, protect the router / ffn_down_exps / shared experts / attention).
Allocation (the surgery)
| tensor role | atom | |
|---|---|---|
cold expert bulk (ffn_gate/up_exps) |
IQ1_KT |
crushed |
ffn_down_exps (residual writer) |
IQ2_KT |
protected |
ffn_gate_inp (router) |
Q6_K |
kept high |
| shared experts | IQ2_KT / IQ3_KT |
protected |
| attention q, output | IQ2_KT |
protected |
| attention k, v | IQ1_KT |
crushed |
| first-2 / last-2 blocks | IQ2_KT |
protected |
| token embeddings / output head | Q4_K / Q6_K |
kept |
Measured notes
PollardMix beats the uniform 1-bit trellis quant on every metric — PPL −4.9%, Mean KLD −14%, Median KLD −20%, top-1 +2.3 pts — at +6.9% size, under the 2-bit ceiling. Same clean sweep as the 7B/14B dense cards, now reproduced on a MoE — the automap policy generalizes (crush cold experts, protect the router / ffn_down_exps / shared experts / attention).
Download a specific file
pip install -U "huggingface_hub[cli]"
hf download PollardWeights/Qwen3-30B-A3B-Pollard \
--include "Qwen3-30B-A3B-Pollard-IQ4_XS.gguf" --local-dir ./
How to run
These are standard GGUF and run with llama.cpp:
llama-server -hf PollardWeights/Qwen3-30B-A3B-Pollard:IQ4_XS
or from a local file:
llama-cli -m Qwen3-30B-A3B-Pollard-IQ4_XS.gguf -ngl 99 -p "Explain why the sky is blue."
llama-server -m Qwen3-30B-A3B-Pollard-IQ4_XS.gguf -ngl 99 # OpenAI-compatible API + web UI at :8080
They also work in anything built on llama.cpp — LM Studio, koboldcpp, Jan, ramalama, Ollama (ollama run hf.co/PollardWeights/Qwen3-30B-A3B-Pollard).
ARM / AVX
llama.cpp repacks weights into an interleaved layout at load time for faster inference on ARM and AVX machines — no special file needed, online repacking covers these quants. The old Q4_0_4_4/4_8/8_8 variants are not required.
Errata
- Trellis (
IQ*_KT) quants need ik_llama.cpp to build/run; K-quants run in any recent llama.cpp. - Measured allocation places bits by per-layer sensitivity under a size budget.
- Single machine; replication invited.
Credits & license
- Base model:
Qwen/Qwen3-30B-A3B - Quantization tooling: llama.cpp (ggml-org)
- Method + tooling: Pollard Weights — measure first, no claim before a number.
- License:
apache-2.0, inherited from the base model.
Built with Pollard Weights — frontier models, small hardware, no compromise.