MiMo-V2.6-Flash-MOPD · NVFP4
309B total · 15B active · text + image + video + audio · 1M context
Xiaomi's own 4-bit experts, re-encoded for Blackwell FP4 tensor cores, with no weight value changed.
Converted by LibertAI · not affiliated with Xiaomi
✨ Why MOPD
MiMo-V2.6-Flash-MOPD is Xiaomi's MOPD2 upgrade of MiMo-V2.6-Flash-RL. Several domain-specialized teachers are distilled into the model on-policy. The headline fix is tool-call repetition, where an agent keeps issuing the same or nearly the same tool call, looks busy, and makes no progress.
Response-level repetition rate, RL-stage vs MOPD, across context lengths and agent harnesses. Figure: Xiaomi MiMo.
If you run MiMo in an agent loop, this is the checkpoint you want. This repo gives you that checkpoint in a format Blackwell GPUs run natively.
🧬 What's inside
Xiaomi ships MiMo-V2.6 already quantized. The routed experts are MXFP4 and the attention is FP8. There is no BF16 checkpoint to start from. Re-quantizing would round values that were already rounded, so we don't quantize anything. We re-encode the expert scales into NVFP4's layout and leave every 4-bit weight code where it is.
flowchart LR
subgraph MX["Released: MXFP4"]
A["32 × E2M1 codes"]
B["1 × E8M0 scale<br/>(power of two)"]
end
subgraph NV["This repo: NVFP4"]
C["32 × E2M1 codes<br/><b>unchanged</b>"]
D["2 × E4M3 scales<br/>(one per 16)"]
E["× per-tensor FP32<br/>weight_scale_2 = 2^k"]
end
A -- "copied bit-for-bit" --> C
B -- "2^e = E4M3(2^(e−k)) · 2^k<br/>exact for every block" --> D
D --- E
Each E8M0 scale is a power of two. E4M3 represents every power of two from 2⁻⁹ to 2⁸ exactly, which is 17 octaves. The widest exponent range inside any MOPD expert is 13 octaves, so each block scale fits exactly with no rounding. gate_proj and up_proj share one global scale, so the fused w13 GEMM stays exact too.
| component | format | tensors |
|---|---|---|
routed experts {gate, up, down}_proj (47 MoE layers × 256) |
NVFP4 · group 16 · W4A4 | 36,096 |
qkv_proj, layer-0 dense MLP, 3 MTP layers |
FP8 block-128 as released | 63 |
o_proj, router, embeddings, lm_head, norms, sinks |
BF16 / F32 as released | 889 total non-expert |
| vision tower (681M) · audio encoder · audio tokenizer | as released | ✓ |
| DFlash drafter sidecar · tokenizer · chat template | as released | ✓ |
Quantization metadata: ModelOpt MIXED_PRECISION in config.json. All multimodal paths are byte-identical to Xiaomi's release.
🎯 Activation scales
W4A4 also quantizes activations to FP4, which needs one global input_scale per expert projection. A placeholder of 1.0 is a known failure mode: at long context, fine-grained block scales underflow. Our checkpoint carries per-layer calibrated values, one for gate/up and one for down in each of the 47 MoE layers, constant across experts. They vary by four orders of magnitude, from ~0.0034 up to 44.4 on layer 47's down_proj.
Provenance. These 94 values come from the calibration in
primitive-ai/MiMo-V2.6-Flash-RL-NVFP4, measured on the RL checkpoint. MOPD is a short distillation run on top of RL, so activation ranges should be close. They were not re-derived on MOPD activations. The expert weights are exact either way. Backends that keep activations in BF16 (W4A16 / Marlin) ignore these scales.
🔬 Verification
1 · Gold test before touching MOPD. Our converter ran on the RL source and was compared byte-for-byte against primitive-ai's published RL NVFP4. weight, weight_scale and weight_scale_2 were identical on every sampled tensor (layers 1 / 5 / 24 / 47), and the generated quantized_layers map matched theirs on all 36,159 entries. The recipe is reproduced exactly.
2 · Independent audit of this repo. A separate checker decodes scales from both the source and the output rather than reusing the converter's math. It ran over all 129 shards:
| check | result |
|---|---|
| expert weight values with identical E2M1 codes | 302,795,194,368 / 302,795,194,368 |
expert blocks where E4M3 × weight_scale_2 ≠ source MX scale |
0 |
| gate/up pairs sharing one global scale | 12,032 / 12,032 |
| non-expert tensors byte-identical to source | 889 / 889 |
| auxiliary files byte-identical (vision, audio, DFlash, tokenizer, template) | all |
| index entries · missing files | 145,273 · 0 |
Source: XiaomiMiMo/MiMo-V2.6-Flash-MOPD @ 2479e2d0029eca9a34cc7e7f55a121925f81908e.
Exact stored weights do not mean bit-identical inference. FP4 activation quantization, kernels and accumulation order all change outputs slightly, just as they do on the released MXFP4 checkpoint.
🚀 Serving with vLLM
Use vLLM's MiMo-V2.6 per-model image. Its MiMo loader expects Xiaomi's FP8 naming (weight_scale_inv, 2-D). ModelOpt's FP8 layers name the parameter weight_scale and store it 4-D. primitive-ai publishes a 21-line loader bridge for this layout, and it applies here unchanged:
# weights + loader bridge
hf download LibertAIDAI/MiMo-V2.6-Flash-MOPD-NVFP4 --local-dir ./MiMo-V2.6-Flash-MOPD-NVFP4
hf download primitive-ai/MiMo-V2.6-Flash-RL-NVFP4 vllm_patch/mimo_v2.py --local-dir ./patch
SP=/usr/local/lib/python3.12/dist-packages/vllm
docker run --gpus all --ipc=host --shm-size 32g -p 8000:8000 \
-v $PWD:/models \
-v $PWD/patch/vllm_patch/mimo_v2.py:$SP/model_executor/models/mimo_v2.py:ro \
vllm/vllm-openai:mimo-v26-x86_64-cu130 \
--model /models/MiMo-V2.6-Flash-MOPD-NVFP4 \
--served-model-name mimo-v2.6-flash-mopd \
--tensor-parallel-size 2 --trust-remote-code --generation-config vllm \
--max-model-len 131072 --gpu-memory-utilization 0.95 \
--reasoning-parser mimo --tool-call-parser mimo --enable-auto-tool-choice
| Hardware | Blackwell (sm_100 / sm_120). 2× 96 GB (RTX PRO 6000) at TP=2 holds 131K context, per the RL build's measurements |
| Without the patch | the load stops with a KeyError on weight_scale_inv |
| Audio off | --limit-mm-per-prompt '{"image":4,"video":0,"audio":0}' skips loading the audio encoder |
| Speculative decoding | the DFlash drafter ships as released, but it was trained against the RL weights, so measure acceptance on MOPD before relying on it |
📈 For throughput and quality numbers on the identical RL layout, see primitive-ai's card, which reports 1.32× the release's throughput at concurrency 16.
🧱 Model at a glance
| Architecture | MiMoV2ForCausalLM: sparse MoE, 309B total / 15B active |
| Layers | 48 (39 sliding-window-128 + 9 global attention), attention sinks |
| Experts | 256 routed, top-8, sigmoid routing (noaux_tc) |
| Heads | 64 Q · 4 KV (global) / 8 KV (SWA) · QK 192 / V 128 |
| MTP | 3 layers in the checkpoint + DFlash drafter sidecar |
| Modalities | text · image · video · audio (681M ViT, 308M audio tokenizer) |
| Context | 1,048,576 tokens |
MiMo-V2.6 architecture. Figure: Xiaomi MiMo.
🙏 Credits & license
- Base model: Xiaomi MiMo, MIT. All capability belongs to them. This repo only changes how the expert scales are stored.
- Recipe, calibration & loader bridge: primitive-ai, whose RL build this reproduces byte-for-byte.
- Conversion & verification: LibertAI. CPU-only, pure numpy, about 5 minutes for 170 GiB.
Released under the MIT license, same as the base model.