ProCreations/MiMo-V2.6-Flash-MOPD-NVFP4

🤗 Hugging Face 来源text-generationmit159B 参数175 GBGGUF✓ 148 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo ProCreations/MiMo-V2.6-Flash-MOPD-NVFP4 ./model-folder
需要做种者 →

MiMo-V2.6-Flash-MOPD: exact NVFP4 weight conversion

This converts the released MXFP4 routed experts of XiaomiMiMo/MiMo-V2.6-Flash-MOPD to NVFP4 without changing any expert weight value. There is no calibration, clipping, optimization or nearest-code rounding. It is a format transcode of the already quantized upstream checkpoint, not a new 4-bit quantization of BF16 weights. Hugging Face's quantized base-model relation describes the stored format; it does not mean an additional lossy quantization was performed here.

MiMo-V2.6-Flash-MOPD is Xiaomi's MOPD upgrade of MiMo-V2.6-Flash-RL. It mitigates tool-call repetition in agent harnesses; see the upstream model card and Xiaomi's blog post. This repository is the same conversion, with the same tools, as ProCreations/MiMo-V2.6-Flash-RL-NVFP4, applied to the MOPD weights.

Source: XiaomiMiMo/MiMo-V2.6-Flash-MOPD, revision 2479e2d0029eca9a34cc7e7f55a121925f81908e. The original model and tokenizer remain subject to the upstream MIT license.

What changed

  • Every E2M1 weight code is unchanged.
  • Each E8M0 scale for 32 weights becomes two exact E4M3 scales for 16 weights, combined with an exact power-of-two global scale per layer. All scale groups were checked for representability; none needed rounding.
  • Dense FP8 tensors are reconstructed in FP32 to preserve their decoded values. Fused attention weights are reordered from the checkpoint's TP=4 interleaving to global Q/K/V order, including its per-rank scale padding.
  • Every other indexed BF16/F32 tensor keeps its bytes. The multimodal encoders, audio tokenizer, tokenizer files and the upstream DFlash sidecars are retained unchanged.
  • ModelOpt mixed-precision metadata selects W4A16_NVFP4 for the experts. This checkpoint's configuration requests no FP4 activation quantization.

The root tensor payload is 193,964,338,432 bytes, excluding auxiliary files.

Verification

verification.json records an independent audit of all 129 root shards:

Check Verified
Expert weight values: identical codes and exactly reconstructed scales 302,795,194,368
Dense FP8 values: independently decoded and reordered into FP32 3,859,808,256
Other tensors: identical bytes 763

Every source shard's SHA-256 was also compared with the pinned Hugging Face LFS manifest (reports/source-manifest.json). auxiliary-verification.json covers the 21 retained auxiliary files.

GGUF. reports/gguf-weight-verification.json checks every expert code and scale in the NVFP4 GGUF against the source checkpoint, with no tolerance, and compares every other tensor with an independent GGUF export of the original MXFP4 checkpoint: 302,795,194,368 expert values and 6,971,406,720 dense values. reports/gguf-release-verification.json confirms that the ten published shards store those bytes, and that the 103 dense matrices stored as BF16 reconstruct their F32 values exactly.

Exact stored weights do not imply bit-identical inference. Runtime dtype, activation quantization, attention kernels, accumulation and sampling can change outputs.

Relation to the RL conversion

  • Same as MiMo-V2.6-Flash-RL: the architecture, configuration, tokenizer, modeling code, audio tokenizer and DFlash draft weights. config.json differs only in the source revision recorded by the conversion, and the chat template differs by one trailing blank line.
  • GGUF: the same 649 tensors with the same names, types and shapes, and the same metadata.
  • What Xiaomi changed (comparison): the expert weight codes, attention, token embeddings and output head.
  • What is identical: the routers and router biases, the MTP layers, and the expert scales of 46 of the 47 MoE layers.
  • The vision encoder, audio encoder and speech embeddings are identical to RL's (459 tensors, SHA-256 equal; RL, MOPD), so an image/audio projector built from either checkpoint holds the same weights.
  • The BF16 DFlash draft GGUF differs from RL's because it embeds the target's token embeddings (with DFlash's trained mask row), so it was rebuilt from MOPD. The draft's own layers are unchanged.

Files

  • Root: the NVFP4 safetensors checkpoint (ModelOpt MIXED_PRECISION, W4A16 NVFP4 experts), config, tokenizer, multimodal encoders and DFlash sidecars.
  • gguf/: the same weights as a 10-shard NVFP4 GGUF for llama.cpp-based runtimes, plus the BF16 DFlash draft GGUF.
  • reports/: verification and workstation measurements.
  • tools/: the conversion and verification scripts. They are the RL release's tools with the source revision and paths as parameters; pipeline.sh runs them in order.

Workstation runtime

The GGUF package was tested on one RTX PRO 6000 (96 GB) with 128 GB of host RAM. It ran with the single-GPU llama.cpp recipe and exactly the RL profile's settings: 524,288-token context, FP8 KV, the DFlash draft, and image and audio input. Measured on the same day and the same runtime (details):

Measurement RL MOPD
Coding decode, 4 tasks (tok/s) 72.3 68.6
DFlash acceptance, coding 0.607 0.572
Decode after a 64K prompt (tok/s) 68.2 70.2
Agent-style turns, decode (tok/s) 56.8 57.4
Cold prefill, 16K / 131K prompt (tok/s) 6,300 / 4,570 6,304 / 4,518

Prefill is unchanged. Decode follows the DFlash draft's acceptance: Xiaomi's draft layers are the same as for RL, and they predict MOPD's coding output slightly less often. The MOPD profile passed the same checks as RL:

  • API and tool compatibility;
  • a repository-review workload;
  • prefix reuse;
  • synthetic retrieval at 131K and 262K tokens;
  • image reading and audio transcription on the production server.

Not tested

  • vLLM or SGLang loading of this safetensors checkpoint. The RL conversion loaded in vLLM 0.29.0 with tensor parallelism 2; this conversion has the same format and configuration but was not re-run there.
  • Near-capacity (512K) retrieval, which failed for the RL model on the same runtime.

Use the upstream recommended sampling: temperature 1.0, top_p 0.95.