MiMo-V2.6 Distill Qwen 9B DWM a1p1 — Q3_K_M GGUF
This repository contains a Q3_K_M GGUF conversion of a DWM-modified
MiMo-V2.6 Distill Qwen 9B checkpoint, along with the BF16 vision projector
required for multimodal use.
Files
| File | Purpose | Size |
|---|---|---|
MiMo-V2.6-Distill-Qwen-9B-DWM-a1p1-Q3_K_M.gguf |
Main Q3_K_M model | 4.62 GB |
mmproj-MiMo-V2.6-Distill-Qwen-9B-DWM-a1p1-BF16.gguf |
BF16 vision projector | 0.92 GB |
The intermediate BF16 main-model GGUF is intentionally not included.
Conversion and validation
- Quantization:
Q3_K_M(4.12 effective bits per weight) - Architecture metadata:
qwen35 - Decoder blocks: 32
- Trained context length: 262,144 tokens
- Main-model tensors: 427
- Vision-projector tensors: 334
- llama.cpp commit:
d1a92352cbd417fd840b4e765c0b82f5fe3d1d89
The source configuration advertised one MTP/NextN layer but contained no
serialized mtp.* tensors. The corrected conversion therefore excluded the
absent NextN layer. The main model and vision projector were load-tested
together, and a short text-generation smoke test completed successfully.
Usage
Use a recent llama.cpp build with MiMo/Qwen3.5 multimodal support:
llama-server \
-m MiMo-V2.6-Distill-Qwen-9B-DWM-a1p1-Q3_K_M.gguf \
--mmproj mmproj-MiMo-V2.6-Distill-Qwen-9B-DWM-a1p1-BF16.gguf \
--ctx-size 262144 \
--jinja
Actual usable context depends on available memory. File hashes are recorded
in SHA256SUMS.
Access
This is a public, manually gated research release. Submit an access request through the Hugging Face repository page.