MiMo-V2.6-Distill-Qwen-9B — GGUF quantizations (XYZ)
GGUF quantizations of MiMo-V2.6-Distill-Qwen-9B (multimodal, hybrid SSM + attention, 32 layers), converted with upstream llama.cpp and cafe-llama.cpp.
What the XYZ series is (and isn't)
XYZ is 100% focused on coding. These quantizations are built and tuned for programming work: the per-tensor recipe keeps the parts that matter for code generation (attention K/V, SSM state) in higher precision, and the build is evaluated so the models stay reliable on code and agentic tasks — even at low bits.
That focus is deliberate. General-knowledge domains are not the goal. There are already plenty of generalist models out there — this one is the coding and agentic one. If you need the best answer on biology, history, politics, or other non-code topics, a general-purpose build of the base model will usually serve you better. XYZ optimizes for coding and agentic behavior first, and that's the trade you're opting into here.
TL;DR — reach for XYZ when the job is writing, reading, or debugging code. For everything else it's still a capable general model, but don't expect it to out-perform an untuned build on pure knowledge tasks.
Models
| File | Size (GiB) | BPW |
|---|---|---|
MiMo-V2.6-Distill-Qwen-9B-Q2.5-XYZ.gguf |
3.29 | 3.15 |
MiMo-V2.6-Distill-Qwen-9B-Q3-XYZ.gguf |
3.62 | 3.47 |
MiMo-V2.6-Distill-Qwen-9B-Q3.5-XYZ.gguf |
4.25 | 4.06 |
MiMo-V2.6-Distill-Qwen-9B-Q4-XYZ.gguf |
4.78 | 4.58 |
MiMo-V2.6-Distill-Qwen-9B-Q4.5-XYZ.gguf |
5.63 | 5.39 |
MiMo-V2.6-Distill-Qwen-9B-Q5-XYZ.gguf |
5.82 | 5.58 |
MiMo-V2.6-Distill-Qwen-9B-Q5.5-XYZ.gguf |
6.18 | 5.92 |
MiMo-V2.6-Distill-Qwen-9B-Q6-XYZ.gguf |
7.01 | 6.73 |
MiMo-V2.6-Distill-Qwen-9B-Q7-XYZ.gguf |
7.24 | 6.95 |
MiMo-V2.6-Distill-Qwen-9B-Q8-XYZ.gguf |
7.84 | 7.52 |
MiMo-V2.6-Distill-Qwen-9B-Q9-XYZ.gguf |
10.70 | 10.27 |
Usage
Serve with llama-server (OpenAI-compatible API):
llama-server -m MiMo-V2.6-Distill-Qwen-9B-Q4-XYZ.gguf --port 8080 -ngl 99 -c 32768
Recommended sampling: --temp 0.6 for balanced, coherent output.
Recommended Settings
llama-server -m MiMo-V2.6-Distill-Qwen-9B-Q4-XYZ.gguf \
--host 0.0.0.0 --port 8080 --ctx-size 128000 -b 2048 --parallel 1 \
-ngl 99 --threads 8 -ub 512 -ctk q8_0 -ctv q8_0 -fa on -kvu \
--temp 0.6 --reasoning-budget 2048 \
--chat-template-kwargs "{\"preserve_thinking\": true, \"reasoning_effort\": \"low\"}" \
--reasoning-budget-message "... I am thinking for too long -- let me gather more info about the task."
Using
--reasoning-budget 2048is like setting thinking to medium, use512for low thinking and8192for high, and omit for uncapped.
Extreme quantizations, sampling to reduce hallucinations
The low-bit files (Q1/Q2) are more prone to hallucination — use this tuned sampling with llama-server for much more stable output:
llama-server -m Qwen3.8-27B-Q2-XYZ-v2.gguf --host 0.0.0.0 --port 1234 \
--temp 0.3 \
--top-p 0.9 \
--top-k 40 \
--repeat-penalty 1.10 \
--repeat-last-n 512 \
--dry-multiplier 0.8 \
--dry-base 1.75 \
--dry-allowed-length 2
Notes
- Model is Apache-2.0, architecture
Qwen3_5ForConditionalGeneration(hybrid SSM + attention, full attention every 4th layer), vocab 248,320,tie_word_embeddings=false. - Quantized with cafe-llama.cpp using a calibrated importance matrix (
imatrix). Key tensors (attention K/V, SSM params) are kept in high precision (BF16 / F32 for norms and conv1d) — full precision where it matters. - Sizes are exact file sizes on disk (GiB, base 1024 — matches the sizes shown on the HF file browser).