pfeifferj/MiMo-V2.6-Flash-RL-GSQ-RCO-GGUF

🤗 Hugging Face 来源text-generationmit251 GBGGUF✓ 3 个校验和今天更新
需要做种者 →

TL;DR: Text-only MiMo-V2.6-Flash-RL at 115.7 GB (3-bit) or 135.0 GB (3.5-bit).

MiMo-V2.6-Flash-RL · GSQ-RCO GGUFs

3-bit and 3.5-bit GGUFs made with GSQ and RCO.

Independent community reproduction. These files were produced by a third party using the published GSQ and RCO methods. They are not an IST-DASLab release and carry no endorsement from the authors of either paper.


Overview

These GGUFs contain the text model, tokenizer and chat template from XiaomiMiMo/MiMo-V2.6-Flash-RL. They support text chat, including thinking mode. Vision, audio and multi-token prediction weights are excluded.

The source uses MXFP4 expert weights, FP8 scaled matrices and BF16/F32 exceptions. The source reference below preserves MXFP4 expert weights and decodes FP8 matrices to BF16.

Method Description
GSQ (Gumbel-Softmax Quantization, paper, code) Post-training scalar quantization that jointly learns per-coordinate grid assignments and per-group scales through a Gumbel-Softmax relaxation.
RCO (Riemannian Constrained Optimization, paper, code) Assigns one of K quantization types to each of N tensors under an exact total size budget, reformulated as a smooth Riemannian manifold in logit space.

Both methods were developed at the Deep Algorithms and Systems Lab (DASLab), Institute of Science and Technology Austria.


Results

MMLU-Pro: no reasoning, 2,048-token context, zero generated tokens. The test uses 2,000 questions across all 14 categories. Zero-shot scoring chooses the answer with the highest log likelihood for its full letter continuation ( A, B, etc.), without a chat template. Question IDs and scores.

Perplexity covers 4,088 next-token predictions across eight 1,024-token contexts of held-out text. Lower is better.

Build Size PPL MMLU-Pro
Source reference 172.932 GB 3.2528 57.70%
GSQ-RCO 3-bit 115.679 GB 3.3754 57.35%
Native imatrix 3-bit 113.519 GB 3.4679 56.20%
Matched initializer 3-bit 115.679 GB 3.3617 57.10%
GSQ-RCO 3.5-bit 134.982 GB 3.3558 57.25%
Native imatrix 3.5-bit 133.653 GB 3.3923 55.05%
Matched initializer 3.5-bit 134.982 GB 3.3540 57.50%

Perplexity rises 3.77% at 3-bit and 3.17% at 3.5-bit over the source reference. MMLU-Pro drops by 0.35 and 0.45 percentage points, respectively.

Matched initializers use the same tensor formats and file sizes as the GSQ builds, before GSQ optimization. They have slightly lower PPL at both sizes; MMLU-Pro changes by +0.25 and -0.25 points after GSQ. Native imatrix controls use separate allocations and smaller files.

GSM8K uses eight questions, thinking enabled and a 2,048-token output cap; scoring checks the numerical answer after ####. IFEval uses 16 prompts, thinking disabled, strict scoring and a 1,024-token cap. Both use temperature 0 and an 8,192-token context.

Build GSM8K IFEval strict IFEval at token cap
Source reference 7/8 15/16 2/16
GSQ-RCO 3-bit 8/8 14/16 2/16
GSQ-RCO 3.5-bit 7/8 14/16 3/16

Counting capped answers as failures gives IFEval scores of 13/16, 12/16 and 11/16, respectively. No GSM8K answer reached the cap.


Usage

llama.cpp

Tested with upstream llama.cpp at 58367713.

hf download pfeifferj/MiMo-V2.6-Flash-RL-GSQ-RCO-GGUF MiMo-V2.6-Flash-RL-GSQ-RCO-3.5bit.gguf --local-dir .

llama-server -m MiMo-V2.6-Flash-RL-GSQ-RCO-3.5bit.gguf \
  -ngl 999 -c 8192 -np 1 -fa on -ctk f32 -ctv f32 --jinja \
  --temp 1.0 --top-p 0.95

These are MiMo's recommended temperature and top-p settings. Thinking is on by default; disable it with --reasoning off.


Quantization procedure

Source: XiaomiMiMo/MiMo-V2.6-Flash-RL at 3b38d063. The text model has 308,778,780,864 stored parameters. Routed expert tensors mix Q2_K, Q3_K and MXFP4. The 3-bit build uses Q8_0 for the 101 non-expert tensors stored as BF16 in the source GGUF; the 3.5-bit build retains BF16. Both retain the 230 F32 tensors.

  1. Capture expert inputs from one million training tokens and 32,768 validation tokens across 47 routed layers. Keep up to 128 training and 64 validation rows per expert.
  2. Run 40 GSQ updates per quantized projection with 32 rows per update. Fit gate, up and down outputs separately; down projections use the original gate/up activations. Native scales and offsets stay fixed. Keep source MXFP4 down candidates unchanged, and retain initializers for poorly sampled experts or when validation worsens.
  3. Run 50 RCO steps per size on four calibration sequences and select the allocation on four separate validation sequences. Assemble the selected packed tensors into GGUF without requantization. Quantization settings.

Citation

This release

@misc{mimov26flashrlgsqrco2026,
  title        = {MiMo-V2.6-Flash-RL GSQ-RCO GGUF quantizations},
  author       = {Josephine Pfeiffer},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/pfeifferj/MiMo-V2.6-Flash-RL-GSQ-RCO-GGUF}}
}

Base model

@misc{mimo2026v26flash,
  title  = {MiMo-V2.6-Flash-RL},
  author = {{Xiaomi MiMo Team}},
  year   = {2026},
  url    = {https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL}
}

Methods

@article{gsq2026,
  title  = {GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling},
  author = {Dadgarnia, Alireza and Tabesh, Soroush and Nikdan, Mahdi and Helcig, Michael and Kurtic, Eldar and Kleinegger, Maximilian and Alistarh, Dan},
  journal= {arXiv preprint arXiv:2604.18556},
  year   = {2026}
}
@article{rco2026,
  title  = {Model Compression with Exact Budget Constraints via Riemannian Manifolds},
  author = {Helcig, Michael and Alistarh, Dan},
  journal= {arXiv preprint arXiv:2605.00649},
  year   = {2026}
}

Acknowledgements

Thanks to Xiaomi MiMo for MiMo-V2.6-Flash-RL and to the Deep Algorithms and Systems Lab (DASLab) for GSQ, RCO and their reference implementations.


License

The weights use the MIT license declared in the source model card. See LICENSE.