quimmedes/MiMo-V2.6-Distill-Qwen-9B-XYZ-GGUF

🤗 Hugging Face sourceapache-2.09B activated71 GBGGUF✓ 11 checksumsupdated today
Needs seeder →

MiMo-V2.6-Distill-Qwen-9B — GGUF quantizations (XYZ)

GGUF quantizations of MiMo-V2.6-Distill-Qwen-9B (multimodal, hybrid SSM + attention, 32 layers), converted with upstream llama.cpp and cafe-llama.cpp.

What the XYZ series is (and isn't)

XYZ is 100% focused on coding. These quantizations are built and tuned for programming work: the per-tensor recipe keeps the parts that matter for code generation (attention K/V, SSM state) in higher precision, and the build is evaluated so the models stay reliable on code and agentic tasks — even at low bits.

That focus is deliberate. General-knowledge domains are not the goal. There are already plenty of generalist models out there — this one is the coding and agentic one. If you need the best answer on biology, history, politics, or other non-code topics, a general-purpose build of the base model will usually serve you better. XYZ optimizes for coding and agentic behavior first, and that's the trade you're opting into here.

TL;DR — reach for XYZ when the job is writing, reading, or debugging code. For everything else it's still a capable general model, but don't expect it to out-perform an untuned build on pure knowledge tasks.

Models

File Size (GiB) BPW
MiMo-V2.6-Distill-Qwen-9B-Q2.5-XYZ.gguf 3.29 3.15
MiMo-V2.6-Distill-Qwen-9B-Q3-XYZ.gguf 3.62 3.47
MiMo-V2.6-Distill-Qwen-9B-Q3.5-XYZ.gguf 4.25 4.06
MiMo-V2.6-Distill-Qwen-9B-Q4-XYZ.gguf 4.78 4.58
MiMo-V2.6-Distill-Qwen-9B-Q4.5-XYZ.gguf 5.63 5.39
MiMo-V2.6-Distill-Qwen-9B-Q5-XYZ.gguf 5.82 5.58
MiMo-V2.6-Distill-Qwen-9B-Q5.5-XYZ.gguf 6.18 5.92
MiMo-V2.6-Distill-Qwen-9B-Q6-XYZ.gguf 7.01 6.73
MiMo-V2.6-Distill-Qwen-9B-Q7-XYZ.gguf 7.24 6.95
MiMo-V2.6-Distill-Qwen-9B-Q8-XYZ.gguf 7.84 7.52
MiMo-V2.6-Distill-Qwen-9B-Q9-XYZ.gguf 10.70 10.27

Usage

Serve with llama-server (OpenAI-compatible API):

llama-server -m MiMo-V2.6-Distill-Qwen-9B-Q4-XYZ.gguf --port 8080 -ngl 99 -c 32768

Recommended sampling: --temp 0.6 for balanced, coherent output.

Recommended Settings

llama-server -m MiMo-V2.6-Distill-Qwen-9B-Q4-XYZ.gguf \
  --host 0.0.0.0 --port 8080 --ctx-size 128000 -b 2048 --parallel 1 \
  -ngl 99 --threads 8 -ub 512 -ctk q8_0 -ctv q8_0 -fa on -kvu \
  --temp 0.6 --reasoning-budget 2048 \
 --chat-template-kwargs "{\"preserve_thinking\": true, \"reasoning_effort\": \"low\"}" \
  --reasoning-budget-message "... I am thinking for too long -- let me gather more info about the task."

Using --reasoning-budget 2048 is like setting thinking to medium, use 512 for low thinking and 8192 for high, and omit for uncapped.

Extreme quantizations, sampling to reduce hallucinations

The low-bit files (Q1/Q2) are more prone to hallucination — use this tuned sampling with llama-server for much more stable output:

llama-server -m Qwen3.8-27B-Q2-XYZ-v2.gguf --host 0.0.0.0 --port 1234 \
  --temp 0.3 \
  --top-p 0.9 \
  --top-k 40 \
  --repeat-penalty 1.10 \
  --repeat-last-n 512 \
  --dry-multiplier 0.8 \
  --dry-base 1.75 \
  --dry-allowed-length 2

Notes

  • Model is Apache-2.0, architecture Qwen3_5ForConditionalGeneration (hybrid SSM + attention, full attention every 4th layer), vocab 248,320, tie_word_embeddings=false.
  • Quantized with cafe-llama.cpp using a calibrated importance matrix (imatrix). Key tensors (attention K/V, SSM params) are kept in high precision (BF16 / F32 for norms and conv1d) — full precision where it matters.
  • Sizes are exact file sizes on disk (GiB, base 1024 — matches the sizes shown on the HF file browser).