MiMo-V2.6-Flash-RL · EXL3
EXL3 quantizations of Xiaomi's MiMo-V2.6-Flash, one folder per bitrate and checkpoint. There are two checkpoints here:
- MOPD (folders ending in
-MOPD), quantized from MiMo-V2.6-Flash-MOPD. This is Xiaomi's newer checkpoint and the one to use for agent and tool-calling work. See What MOPD is. - RL (plain bitrate folders), quantized from the original MiMo-V2.6-Flash-RL.
Every pack is scored against its own source checkpoint on held-out text before it is uploaded.
Serving recipe: vcruz305/MiMo-V2.6-Flash-RL-EXL3-recipe. It has the runtime, the DFlash drafter fix, the exact serve command, and the measured decode, prefill and TTFT numbers for both the native /v1 server and a TabbyAPI route. The MOPD packs use the same architecture and load the same way; point the serve command at a -MOPD folder.
| checkpoint | bpw | size | smallest card | top-1 vs original | mean KLD | p99 KLD | download |
|---|---|---|---|---|---|---|---|
| MOPD | 2.20 | 86.94 GB | 96 GB | scores being added | 2.20bpw-MOPD | ||
| MOPD | 2.50 | 98.47 GB | 128 GB | 72.21% (7,394 / 10,240) | 0.5641 | 6.332 | 2.50bpw-MOPD |
| RL | 2.50 | 98.48 GB | 128 GB | 83.76% (8,577 / 10,240) | 0.2055 | 3.380 | 2.50bpw |
| RL | 2.20 | 86.94 GB | 96 GB | 82.16% (8,413 / 10,240) | 0.1926 | 2.241 | 2.20bpw |
- top-1: positions where the pack's most likely next token matches the reference's, counted over 10,240 positions of held-out text.
- KLD: KL(reference ‖ pack) per position over the same positions, mean and 99th percentile. Lower is closer.
- Reference: Xiaomi's own implementation (
modeling_mimo_v2.py) running the original weights with fp32 activations, on the same tokens. Each pack is compared with the checkpoint it was made from: MOPD packs against MOPD, RL packs against RL. - For scale: the unquantized RL weights running in the same exllamav3 code match the reference on 96.49% of positions (9,881 / 10,240) with mean KLD 0.0070. That is about the best any pack here can score.
- On a second set of held-out text (a separate 10,240 positions, also untouched by the quants): RL 2.50 87.30% (KLD 0.1086), RL 2.20 82.64% (KLD 0.1986). MOPD 2.50 77.26% (KLD 0.3796) on that same second set.
- size is the whole folder on disk. smallest card is the smallest GPU that holds the pack resident with room for KV cache. The 2.20 packs fit a single 96 GB card, which is what the serving numbers below were measured on.
What MOPD is
MiMo-V2.6-Flash-MOPD is Xiaomi's upgrade of the Flash-RL checkpoint. It is the same model: same architecture, same size, same config.json, and the same drafter weights. Only the trained weights differ. Xiaomi released it as a separate repository, and Flash-RL itself was not changed.
The update adds one more training stage on top of the RL model, MOPD2 (multi-teacher on-policy distillation). Several domain-specialized teacher models are distilled into the model on its own outputs:
- Verifiable-task teachers (trained with RL) grade full autonomous rollouts.
- Teacher-prefix turns: the model continues from a teacher's conversation history and the teacher scores that turn.
- Demonstration-prefix turns: the model continues from a synthetic demonstration, which covers open-ended domains where a reliable reward is hard to write, such as long-horizon game development, scientific research and embodied tasks.
The practical change is a fix for tool-call repetition. After the V2.6 release, the most noticeable problem in agent use was the model sometimes issuing the same or nearly the same tool call over and over. Nothing errored, but it used up time and context without making progress. The MOPD checkpoint cuts that repetition across context lengths and agent harnesses.
If you run MiMo as an agent or with tools, use the MOPD packs. The RL packs stay here unchanged for anyone who has already deployed or evaluated them.
Details: Xiaomi's MOPD model card, technical report §5.6, and the tool-call repetition write-up.
About these quants
- Made with SAGE, our own mixed-precision quantization method for EXL3. The number in each folder name is the average bitrate of the model body. The output head is 6-bit in every pack.
- The MOPD packs were encoded from the MOPD weights at the same bitrates as the RL packs. They are not the RL packs with a patch applied.
- The scores compare each pack with the reference above on text that was not used to make the quants; no evaluation text was used to calibrate them. They measure how closely a pack tracks the original weights, not task accuracy.
- These packs cover the text model: the 48-layer MoE backbone with its hybrid attention. MiMo's vision and audio encoders and its multi-token-prediction drafter are not part of the pack and are not implemented in the loader, so it serves text.
- Each folder carries the original tokenizer and chat template.
What MiMo-V2.6-Flash is
Before quantization this is a sparse MoE language model from Xiaomi's MiMo team: 309B parameters in total and about 15B active per token, a 1M-token context, and one model covering text, image, video and audio. It is the efficiency-balanced checkpoint of the MiMo-V2.6 series, RL-trained in one mixed run across coding, agents, visual tasks and cybersecurity, with groupwise agentic grading and a multi-teacher distillation stage afterwards. The MOPD checkpoint adds the distillation stage described above. It is MIT licensed.
Architecture
From the model's config.json and Xiaomi's card. Identical for RL and MOPD:
| Layers | 48 (39 sliding-window, 9 full attention) |
| Hidden size | 4,096 |
| Attention | 64 query heads; 8 KV heads on the sliding layers, 4 on the full layers; QK head dim 192, V head dim 128; RoPE, 1M context |
| Sliding window | 128 tokens; attention_chunk 128; learned sink bias on the sliding layers |
| MoE | 256 routed experts per layer, 8 active per token, no shared expert, sigmoid router with normalized weights; first block dense, experts 303B of the model |
| Vocabulary | 152,576 tokens, untied input/output embeddings |
| Storage | experts in MXFP4, attention in block-FP8 (128x128), the rest BF16 |
Reported results
Scores for the original MiMo-V2.6 Flash model from Xiaomi's card, not re-run on these quants. Use the top-1/KLD columns above to see how closely each pack tracks its source.
| Benchmark | Category | MiMo-V2.6 Flash |
|---|---|---|
| DeepSWE v1.1 | Code agent | 67.9 |
| MiMo Code Bench | Code agent | 61.2 |
| AutomationBench v1.0.6 | General agent | 52.3 |
| Toolathlon-Verified | General agent | 73.6 |
| Terminal Bench 2.1 | General agent | 87.6 |
| OSWorld-Verified | General agent | 80.8 |
| JobBench | General agent | 61.2 |
| CyberGym | Cybersecurity | 95.1 |
| SEC Bench Pro | Cybersecurity | 47.5 |
| MiMo VisualCoding | Visual agent | 71.5 |
Using it well
Xiaomi's recommended settings, which apply to these packs as well:
- Sampling:
temperature=1.0,top_p=0.95. - Reasoning: the model reports its thinking through a reasoning parser on the serving side; the chat template ships in each folder.
Running
These packs load in exllamav3. Upstream does not carry the MiMo-V2 architecture yet; the port is turboderp-org/exllamav3#399. Build that branch (it needs the CUDA extension compiled for your card), then:
from exllamav3 import Cache, Config, Model, Tokenizer
config = Config.from_directory("MiMo-V2.6-Flash-RL-EXL3/2.20bpw-MOPD")
model = Model.from_config(config)
cache = Cache(model, max_num_tokens = 32768)
tok = Tokenizer(config)
model.load()
To download only one pack:
hf download vcruz305/MiMo-V2.6-Flash-RL-EXL3 --include "2.20bpw-MOPD/*" --local-dir MiMo-V2.6-Flash-RL-EXL3
examples/chat.py in that repository is an interactive client for the same API:
python examples/chat.py -m MiMo-V2.6-Flash-RL-EXL3/2.20bpw-MOPD -cs 32768
For an OpenAI-compatible endpoint, any exllamav3 front end will do. The serving recipe ships one, with a preflight check and the drafter wired in. The numbers below were measured with a small native /v1 wrapper over the same library, at a 65,536-token cache and up to 16 concurrent requests.
Speeds on an RTX PRO 6000 (96 GB)
The RL 2.20 pack served on one RTX PRO 6000 with the exllamav3 build above, measured by SixCat in 6 minutes of continuous traffic. The MOPD 2.20 pack has the same architecture, bitrate and size, so it serves at the same speed; it has not been re-measured separately.
| Decode, one request | 49.6 tok/s (p95 49.6) |
| Prefill | 2,370 tok/s |
| Balanced serving | 116.2 tok/s aggregate at concurrency 8 |
| Time to first token | 4.68 s p50 (4.71 s p95) |
| Max usable concurrency | 8 |
The weights are resident in VRAM with the KV cache alongside them, so these are the numbers to expect on a card this size and nothing was offloaded.
License and credit
MIT, same as the base models. The models, their training and their evaluation are Xiaomi's work: see the original model cards for Flash-RL and Flash-MOPD, and the technical report, for full details. Quantization is ours.
@misc{mimo2026v26flash,
title={MiMo-V2.6-Flash-RL},
author={{Xiaomi MiMo Team}},
year={2026},
howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL}},
}
@misc{mimo2026v26flashmopd,
title={MiMo-V2.6-Flash-MOPD},
author={{Xiaomi MiMo Team}},
year={2026},
howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-MOPD}},
}