LibertAIDAI/MiMo-V2.6-Flash-MOPD-NVFP4

🤗 Hugging Face sourceimage-text-to-textmit159B params163 GBsafetensors✓ 136 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo LibertAIDAI/MiMo-V2.6-Flash-MOPD-NVFP4 ./model-folder
Needs a seeder →

MiMo-V2.6-Flash-MOPD · NVFP4

309B total · 15B active · text + image + video + audio · 1M context

Xiaomi's own 4-bit experts, re-encoded for Blackwell FP4 tensor cores, with no weight value changed.

Converted by LibertAI · not affiliated with Xiaomi


✨ Why MOPD

MiMo-V2.6-Flash-MOPD is Xiaomi's MOPD2 upgrade of MiMo-V2.6-Flash-RL. Several domain-specialized teachers are distilled into the model on-policy. The headline fix is tool-call repetition, where an agent keeps issuing the same or nearly the same tool call, looks busy, and makes no progress.


Response-level repetition rate, RL-stage vs MOPD, across context lengths and agent harnesses. Figure: Xiaomi MiMo.

If you run MiMo in an agent loop, this is the checkpoint you want. This repo gives you that checkpoint in a format Blackwell GPUs run natively.


🧬 What's inside

Xiaomi ships MiMo-V2.6 already quantized. The routed experts are MXFP4 and the attention is FP8. There is no BF16 checkpoint to start from. Re-quantizing would round values that were already rounded, so we don't quantize anything. We re-encode the expert scales into NVFP4's layout and leave every 4-bit weight code where it is.

flowchart LR
    subgraph MX["Released: MXFP4"]
        A["32 × E2M1 codes"]
        B["1 × E8M0 scale<br/>(power of two)"]
    end
    subgraph NV["This repo: NVFP4"]
        C["32 × E2M1 codes<br/><b>unchanged</b>"]
        D["2 × E4M3 scales<br/>(one per 16)"]
        E["× per-tensor FP32<br/>weight_scale_2 = 2^k"]
    end
    A -- "copied bit-for-bit" --> C
    B -- "2^e = E4M3(2^(e−k)) · 2^k<br/>exact for every block" --> D
    D --- E

Each E8M0 scale is a power of two. E4M3 represents every power of two from 2⁻⁹ to 2⁸ exactly, which is 17 octaves. The widest exponent range inside any MOPD expert is 13 octaves, so each block scale fits exactly with no rounding. gate_proj and up_proj share one global scale, so the fused w13 GEMM stays exact too.

component format tensors
routed experts {gate, up, down}_proj (47 MoE layers × 256) NVFP4 · group 16 · W4A4 36,096
qkv_proj, layer-0 dense MLP, 3 MTP layers FP8 block-128 as released 63
o_proj, router, embeddings, lm_head, norms, sinks BF16 / F32 as released 889 total non-expert
vision tower (681M) · audio encoder · audio tokenizer as released ✓
DFlash drafter sidecar · tokenizer · chat template as released ✓

Quantization metadata: ModelOpt MIXED_PRECISION in config.json. All multimodal paths are byte-identical to Xiaomi's release.


🎯 Activation scales

W4A4 also quantizes activations to FP4, which needs one global input_scale per expert projection. A placeholder of 1.0 is a known failure mode: at long context, fine-grained block scales underflow. Our checkpoint carries per-layer calibrated values, one for gate/up and one for down in each of the 47 MoE layers, constant across experts. They vary by four orders of magnitude, from ~0.0034 up to 44.4 on layer 47's down_proj.

Provenance. These 94 values come from the calibration in primitive-ai/MiMo-V2.6-Flash-RL-NVFP4, measured on the RL checkpoint. MOPD is a short distillation run on top of RL, so activation ranges should be close. They were not re-derived on MOPD activations. The expert weights are exact either way. Backends that keep activations in BF16 (W4A16 / Marlin) ignore these scales.


🔬 Verification

1 · Gold test before touching MOPD. Our converter ran on the RL source and was compared byte-for-byte against primitive-ai's published RL NVFP4. weight, weight_scale and weight_scale_2 were identical on every sampled tensor (layers 1 / 5 / 24 / 47), and the generated quantized_layers map matched theirs on all 36,159 entries. The recipe is reproduced exactly.

2 · Independent audit of this repo. A separate checker decodes scales from both the source and the output rather than reusing the converter's math. It ran over all 129 shards:

check result
expert weight values with identical E2M1 codes 302,795,194,368 / 302,795,194,368
expert blocks where E4M3 × weight_scale_2 ≠ source MX scale 0
gate/up pairs sharing one global scale 12,032 / 12,032
non-expert tensors byte-identical to source 889 / 889
auxiliary files byte-identical (vision, audio, DFlash, tokenizer, template) all
index entries · missing files 145,273 · 0

Source: XiaomiMiMo/MiMo-V2.6-Flash-MOPD @ 2479e2d0029eca9a34cc7e7f55a121925f81908e.

Exact stored weights do not mean bit-identical inference. FP4 activation quantization, kernels and accumulation order all change outputs slightly, just as they do on the released MXFP4 checkpoint.


🚀 Serving with vLLM

Use vLLM's MiMo-V2.6 per-model image. Its MiMo loader expects Xiaomi's FP8 naming (weight_scale_inv, 2-D). ModelOpt's FP8 layers name the parameter weight_scale and store it 4-D. primitive-ai publishes a 21-line loader bridge for this layout, and it applies here unchanged:

# weights + loader bridge
hf download LibertAIDAI/MiMo-V2.6-Flash-MOPD-NVFP4 --local-dir ./MiMo-V2.6-Flash-MOPD-NVFP4
hf download primitive-ai/MiMo-V2.6-Flash-RL-NVFP4 vllm_patch/mimo_v2.py --local-dir ./patch

SP=/usr/local/lib/python3.12/dist-packages/vllm
docker run --gpus all --ipc=host --shm-size 32g -p 8000:8000 \
  -v $PWD:/models \
  -v $PWD/patch/vllm_patch/mimo_v2.py:$SP/model_executor/models/mimo_v2.py:ro \
  vllm/vllm-openai:mimo-v26-x86_64-cu130 \
  --model /models/MiMo-V2.6-Flash-MOPD-NVFP4 \
  --served-model-name mimo-v2.6-flash-mopd \
  --tensor-parallel-size 2 --trust-remote-code --generation-config vllm \
  --max-model-len 131072 --gpu-memory-utilization 0.95 \
  --reasoning-parser mimo --tool-call-parser mimo --enable-auto-tool-choice
Hardware Blackwell (sm_100 / sm_120). 2× 96 GB (RTX PRO 6000) at TP=2 holds 131K context, per the RL build's measurements
Without the patch the load stops with a KeyError on weight_scale_inv
Audio off --limit-mm-per-prompt '{"image":4,"video":0,"audio":0}' skips loading the audio encoder
Speculative decoding the DFlash drafter ships as released, but it was trained against the RL weights, so measure acceptance on MOPD before relying on it

📈 For throughput and quality numbers on the identical RL layout, see primitive-ai's card, which reports 1.32× the release's throughput at concurrency 16.


🧱 Model at a glance

Architecture MiMoV2ForCausalLM: sparse MoE, 309B total / 15B active
Layers 48 (39 sliding-window-128 + 9 global attention), attention sinks
Experts 256 routed, top-8, sigmoid routing (noaux_tc)
Heads 64 Q · 4 KV (global) / 8 KV (SWA) · QK 192 / V 128
MTP 3 layers in the checkpoint + DFlash drafter sidecar
Modalities text · image · video · audio (681M ViT, 308M audio tokenizer)
Context 1,048,576 tokens

MiMo-V2.6 architecture. Figure: Xiaomi MiMo.

🙏 Credits & license

  • Base model: Xiaomi MiMo, MIT. All capability belongs to them. This repo only changes how the expert scales are stored.
  • Recipe, calibration & loader bridge: primitive-ai, whose RL build this reproduces byte-for-byte.
  • Conversion & verification: LibertAI. CPU-only, pure numpy, about 5 minutes for 170 GiB.

Released under the MIT license, same as the base model.