LibertAIDAI/MiMo-V2.6-Flash-MOPD-NVFP4

🤗 Hugging Face 来源image-text-to-textmit159B 参数163 GBsafetensors✓ 136 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo LibertAIDAI/MiMo-V2.6-Flash-MOPD-NVFP4 ./model-folder
需要做种者 →

MiMo-V2.6-Flash-MOPD · NVFP4

309B total · 15B active · text + image + video + audio · 1M context

Xiaomi's own 4-bit experts, re-encoded for Blackwell FP4 tensor cores, with no weight value changed.

Converted by LibertAI · not affiliated with Xiaomi


✨ Why MOPD

MiMo-V2.6-Flash-MOPD is Xiaomi's MOPD2 upgrade of MiMo-V2.6-Flash-RL. Several domain-specialized teachers are distilled into the model on-policy. The headline fix is tool-call repetition, where an agent keeps issuing the same or nearly the same tool call, looks busy, and makes no progress.


Response-level repetition rate, RL-stage vs MOPD, across context lengths and agent harnesses. Figure: Xiaomi MiMo.

If you run MiMo in an agent loop, this is the checkpoint you want. This repo gives you that checkpoint in a format Blackwell GPUs run natively.


🧬 What's inside

Xiaomi ships MiMo-V2.6 already quantized. The routed experts are MXFP4 and the attention is FP8. There is no BF16 checkpoint to start from. Re-quantizing would round values that were already rounded, so we don't quantize anything. We re-encode the expert scales into NVFP4's layout and leave every 4-bit weight code where it is.

flowchart LR
    subgraph MX["Released: MXFP4"]
        A["32 × E2M1 codes"]
        B["1 × E8M0 scale<br/>(power of two)"]
    end
    subgraph NV["This repo: NVFP4"]
        C["32 × E2M1 codes<br/><b>unchanged</b>"]
        D["2 × E4M3 scales<br/>(one per 16)"]
        E["× per-tensor FP32<br/>weight_scale_2 = 2^k"]
    end
    A -- "copied bit-for-bit" --> C
    B -- "2^e = E4M3(2^(e−k)) · 2^k<br/>exact for every block" --> D
    D --- E

Each E8M0 scale is a power of two. E4M3 represents every power of two from 2⁻⁹ to 2⁸ exactly, which is 17 octaves. The widest exponent range inside any MOPD expert is 13 octaves, so each block scale fits exactly with no rounding. gate_proj and up_proj share one global scale, so the fused w13 GEMM stays exact too.

component format tensors
routed experts {gate, up, down}_proj (47 MoE layers × 256) NVFP4 · group 16 · W4A4 36,096
qkv_proj, layer-0 dense MLP, 3 MTP layers FP8 block-128 as released 63
o_proj, router, embeddings, lm_head, norms, sinks BF16 / F32 as released 889 total non-expert
vision tower (681M) · audio encoder · audio tokenizer as released ✓
DFlash drafter sidecar · tokenizer · chat template as released ✓

Quantization metadata: ModelOpt MIXED_PRECISION in config.json. All multimodal paths are byte-identical to Xiaomi's release.


🎯 Activation scales

W4A4 also quantizes activations to FP4, which needs one global input_scale per expert projection. A placeholder of 1.0 is a known failure mode: at long context, fine-grained block scales underflow. Our checkpoint carries per-layer calibrated values, one for gate/up and one for down in each of the 47 MoE layers, constant across experts. They vary by four orders of magnitude, from ~0.0034 up to 44.4 on layer 47's down_proj.

Provenance. These 94 values come from the calibration in primitive-ai/MiMo-V2.6-Flash-RL-NVFP4, measured on the RL checkpoint. MOPD is a short distillation run on top of RL, so activation ranges should be close. They were not re-derived on MOPD activations. The expert weights are exact either way. Backends that keep activations in BF16 (W4A16 / Marlin) ignore these scales.


🔬 Verification

1 · Gold test before touching MOPD. Our converter ran on the RL source and was compared byte-for-byte against primitive-ai's published RL NVFP4. weight, weight_scale and weight_scale_2 were identical on every sampled tensor (layers 1 / 5 / 24 / 47), and the generated quantized_layers map matched theirs on all 36,159 entries. The recipe is reproduced exactly.

2 · Independent audit of this repo. A separate checker decodes scales from both the source and the output rather than reusing the converter's math. It ran over all 129 shards:

check result
expert weight values with identical E2M1 codes 302,795,194,368 / 302,795,194,368
expert blocks where E4M3 × weight_scale_2 ≠ source MX scale 0
gate/up pairs sharing one global scale 12,032 / 12,032
non-expert tensors byte-identical to source 889 / 889
auxiliary files byte-identical (vision, audio, DFlash, tokenizer, template) all
index entries · missing files 145,273 · 0

Source: XiaomiMiMo/MiMo-V2.6-Flash-MOPD @ 2479e2d0029eca9a34cc7e7f55a121925f81908e.

Exact stored weights do not mean bit-identical inference. FP4 activation quantization, kernels and accumulation order all change outputs slightly, just as they do on the released MXFP4 checkpoint.


🚀 Serving with vLLM

Use vLLM's MiMo-V2.6 per-model image. Its MiMo loader expects Xiaomi's FP8 naming (weight_scale_inv, 2-D). ModelOpt's FP8 layers name the parameter weight_scale and store it 4-D. primitive-ai publishes a 21-line loader bridge for this layout, and it applies here unchanged:

# weights + loader bridge
hf download LibertAIDAI/MiMo-V2.6-Flash-MOPD-NVFP4 --local-dir ./MiMo-V2.6-Flash-MOPD-NVFP4
hf download primitive-ai/MiMo-V2.6-Flash-RL-NVFP4 vllm_patch/mimo_v2.py --local-dir ./patch

SP=/usr/local/lib/python3.12/dist-packages/vllm
docker run --gpus all --ipc=host --shm-size 32g -p 8000:8000 \
  -v $PWD:/models \
  -v $PWD/patch/vllm_patch/mimo_v2.py:$SP/model_executor/models/mimo_v2.py:ro \
  vllm/vllm-openai:mimo-v26-x86_64-cu130 \
  --model /models/MiMo-V2.6-Flash-MOPD-NVFP4 \
  --served-model-name mimo-v2.6-flash-mopd \
  --tensor-parallel-size 2 --trust-remote-code --generation-config vllm \
  --max-model-len 131072 --gpu-memory-utilization 0.95 \
  --reasoning-parser mimo --tool-call-parser mimo --enable-auto-tool-choice
Hardware Blackwell (sm_100 / sm_120). 2× 96 GB (RTX PRO 6000) at TP=2 holds 131K context, per the RL build's measurements
Without the patch the load stops with a KeyError on weight_scale_inv
Audio off --limit-mm-per-prompt '{"image":4,"video":0,"audio":0}' skips loading the audio encoder
Speculative decoding the DFlash drafter ships as released, but it was trained against the RL weights, so measure acceptance on MOPD before relying on it

📈 For throughput and quality numbers on the identical RL layout, see primitive-ai's card, which reports 1.32× the release's throughput at concurrency 16.


🧱 Model at a glance

Architecture MiMoV2ForCausalLM: sparse MoE, 309B total / 15B active
Layers 48 (39 sliding-window-128 + 9 global attention), attention sinks
Experts 256 routed, top-8, sigmoid routing (noaux_tc)
Heads 64 Q · 4 KV (global) / 8 KV (SWA) · QK 192 / V 128
MTP 3 layers in the checkpoint + DFlash drafter sidecar
Modalities text · image · video · audio (681M ViT, 308M audio tokenizer)
Context 1,048,576 tokens

MiMo-V2.6 architecture. Figure: Xiaomi MiMo.

🙏 Credits & license

  • Base model: Xiaomi MiMo, MIT. All capability belongs to them. This repo only changes how the expert scales are stored.
  • Recipe, calibration & loader bridge: primitive-ai, whose RL build this reproduces byte-for-byte.
  • Conversion & verification: LibertAI. CPU-only, pure numpy, about 5 minutes for 170 GiB.

Released under the MIT license, same as the base model.