MiMo-V2.6-Flash-MOPD: exact NVFP4 weight conversion
This converts the released MXFP4 routed experts of
XiaomiMiMo/MiMo-V2.6-Flash-MOPD to NVFP4
without changing any expert weight value. There is no calibration, clipping, optimization or
nearest-code rounding. It is a format transcode of the already quantized upstream checkpoint, not a new
4-bit quantization of BF16 weights. Hugging Face's quantized base-model relation describes the stored
format; it does not mean an additional lossy quantization was performed here.
MiMo-V2.6-Flash-MOPD is Xiaomi's MOPD upgrade of MiMo-V2.6-Flash-RL. It mitigates tool-call repetition in agent harnesses; see the upstream model card and Xiaomi's blog post. This repository is the same conversion, with the same tools, as ProCreations/MiMo-V2.6-Flash-RL-NVFP4, applied to the MOPD weights.
Source: XiaomiMiMo/MiMo-V2.6-Flash-MOPD, revision
2479e2d0029eca9a34cc7e7f55a121925f81908e.
The original model and tokenizer remain subject to the upstream MIT license.
What changed
- Every E2M1 weight code is unchanged.
- Each E8M0 scale for 32 weights becomes two exact E4M3 scales for 16 weights, combined with an exact power-of-two global scale per layer. All scale groups were checked for representability; none needed rounding.
- Dense FP8 tensors are reconstructed in FP32 to preserve their decoded values. Fused attention weights are reordered from the checkpoint's TP=4 interleaving to global Q/K/V order, including its per-rank scale padding.
- Every other indexed BF16/F32 tensor keeps its bytes. The multimodal encoders, audio tokenizer, tokenizer files and the upstream DFlash sidecars are retained unchanged.
- ModelOpt mixed-precision metadata selects
W4A16_NVFP4for the experts. This checkpoint's configuration requests no FP4 activation quantization.
The root tensor payload is 193,964,338,432 bytes, excluding auxiliary files.
Verification
verification.json records an independent audit of all 129 root shards:
| Check | Verified |
|---|---|
| Expert weight values: identical codes and exactly reconstructed scales | 302,795,194,368 |
| Dense FP8 values: independently decoded and reordered into FP32 | 3,859,808,256 |
| Other tensors: identical bytes | 763 |
Every source shard's SHA-256 was also compared with the pinned Hugging Face LFS manifest
(reports/source-manifest.json).
auxiliary-verification.json covers the 21 retained auxiliary files.
GGUF. reports/gguf-weight-verification.json checks every expert
code and scale in the NVFP4 GGUF against the source checkpoint, with no tolerance, and compares every other
tensor with an independent GGUF export of the original MXFP4 checkpoint:
302,795,194,368 expert values and 6,971,406,720 dense values.
reports/gguf-release-verification.json confirms that the ten
published shards store those bytes, and that the 103 dense matrices stored as BF16 reconstruct their
F32 values exactly.
Exact stored weights do not imply bit-identical inference. Runtime dtype, activation quantization, attention kernels, accumulation and sampling can change outputs.
Relation to the RL conversion
- Same as MiMo-V2.6-Flash-RL: the architecture, configuration, tokenizer, modeling code, audio tokenizer and DFlash
draft weights.
config.jsondiffers only in the source revision recorded by the conversion, and the chat template differs by one trailing blank line. - GGUF: the same 649 tensors with the same names, types and shapes, and the same metadata.
- What Xiaomi changed (comparison): the expert weight codes, attention, token embeddings and output head.
- What is identical: the routers and router biases, the MTP layers, and the expert scales of 46 of the 47 MoE layers.
- The vision encoder, audio encoder and speech embeddings are identical to RL's (459 tensors, SHA-256 equal; RL, MOPD), so an image/audio projector built from either checkpoint holds the same weights.
- The BF16 DFlash draft GGUF differs from RL's because it embeds the target's token embeddings (with DFlash's trained mask row), so it was rebuilt from MOPD. The draft's own layers are unchanged.
Files
- Root: the NVFP4 safetensors checkpoint (ModelOpt
MIXED_PRECISION, W4A16 NVFP4 experts), config, tokenizer, multimodal encoders and DFlash sidecars. gguf/: the same weights as a 10-shard NVFP4 GGUF for llama.cpp-based runtimes, plus the BF16 DFlash draft GGUF.reports/: verification and workstation measurements.tools/: the conversion and verification scripts. They are the RL release's tools with the source revision and paths as parameters;pipeline.shruns them in order.
Workstation runtime
The GGUF package was tested on one RTX PRO 6000 (96 GB) with 128 GB of host RAM. It ran with the single-GPU llama.cpp recipe and exactly the RL profile's settings: 524,288-token context, FP8 KV, the DFlash draft, and image and audio input. Measured on the same day and the same runtime (details):
| Measurement | RL | MOPD |
|---|---|---|
| Coding decode, 4 tasks (tok/s) | 72.3 | 68.6 |
| DFlash acceptance, coding | 0.607 | 0.572 |
| Decode after a 64K prompt (tok/s) | 68.2 | 70.2 |
| Agent-style turns, decode (tok/s) | 56.8 | 57.4 |
| Cold prefill, 16K / 131K prompt (tok/s) | 6,300 / 4,570 | 6,304 / 4,518 |
Prefill is unchanged. Decode follows the DFlash draft's acceptance: Xiaomi's draft layers are the same as for RL, and they predict MOPD's coding output slightly less often. The MOPD profile passed the same checks as RL:
- API and tool compatibility;
- a repository-review workload;
- prefix reuse;
- synthetic retrieval at 131K and 262K tokens;
- image reading and audio transcription on the production server.
Not tested
- vLLM or SGLang loading of this safetensors checkpoint. The RL conversion loaded in vLLM 0.29.0 with tensor parallelism 2; this conversion has the same format and configuration but was not re-run there.
- Near-capacity (512K) retrieval, which failed for the RL model on the same runtime.
Use the upstream recommended sampling: temperature 1.0, top_p 0.95.