vcruz305/MiMo-V2.6-Flash-RL-EXL3-2.20bpw-derisked

🤗 Hugging Face 来源text-generationmit43.4B 参数87 GBsafetensors✓ 17 个校验和今天更新
需要做种者 →

MiMo-V2.6-Flash-RL · EXL3 2.20 bpw, derisked

A collaboration between Victor Cruz and Blackfrost (@Blackfrost_AI).

The EXL3 quantization is mine: 2.20 bpw, and it fits and serves on a single RTX PRO 6000 (96 GB). The derisking overlay is Blackfrost's — the same one that ships in Blackfrost-AI/MiMo-v2.6-Flash-RL-Derisked for SGLang — ported here onto the EXL3 serving path. The weights are the 2.20 bpw pack unchanged. The overlay is a runtime edit plus the chat template baked into this repo, not a change to the checkpoint.

Serving recipe: vcruz305/MiMo-V2.6-Flash-RL-EXL3-recipe — the runtime, the DFlash drafter fix, the exact serve command and the measured decode / prefill / TTFT numbers.

Quality against the original checkpoint

Scored against Xiaomi's own implementation running the original weights with fp32 activations, on 10,240 positions of held-out text that was not used to make the quant. The overlay does not touch the weights, so these numbers are the pack's.

top-1 vs original 82.16% (8,413 / 10,240)
mean KLD 0.1926
p99 KLD 2.241
second held-out set 82.64% top-1, mean KLD 0.1986
  • top-1: positions where the pack's most likely next token matches the reference's.
  • KLD: KL(reference ‖ pack) per position, mean and 99th percentile. Lower is closer.
  • For scale: the unquantized weights running in the same exllamav3 code match the reference on 96.49% of positions (9,881 / 10,240) with mean KLD 0.0070. That is about the best any pack can score.
  • The average bitrate of the model body is 2.20; the output head is 6-bit. quantization_config.bits in config.json is rounded to an integer because the Hub's metadata schema rejects a fractional value — the per-tensor bitrates are in quantization_config.json.

The overlay

The directions, the manifest and the hook are under derisk/ and exllamav3/, and the method is Blackfrost's. It applies at serving time and does not modify weights or routing.

It is a load-time switch, not a per-request flag. Same weights, same server, restart to change it.

bash launch.sh <model-dir>            # derisked
bash launch.sh --off <model-dir>      # plain EXL3 2.20

launch.sh sets BLACKFROST_MIMO_MOE_INTERVENTION to derisk/intervention.json. With the variable unset the hook is a no-op, so one patched fork serves both modes.

The hook needs one file and one line added to the exllamav3 fork (vcruz305/exllamav3, branch feat/mimo-v2, commit 93e58ca):

bash setup.sh <path-to-exllamav3-checkout>

setup.sh copies exllamav3/blackfrost_derisk.py into the fork and applies exllamav3/transformer-derisk.patch. It checks the fork commit and refuses to patch anything it doesn't match.

Measured

On one RTX PRO 6000, 655,360 context, Q4 KV cache, with the EXL3 4.0 bpw dflash drafter. The reference column is Blackfrost's own MXFP4 build of the same overlay, scored by the same judge on the same 32-prompt set. It is a higher-bitrate quant, so the gap to it is mostly the base quant, not the overlay.

Harmful set:

actionable actionability
this pack, overlay off 13/32 1.36
this pack, overlay on 26/32 2.38
Blackfrost MXFP4 reference 30/32 2.84

Decode speed, single stream, overlay on vs off: 182.8 vs 184.7 tok/s p50 — about 1-3% depending on the prompt. Draft acceptance was 0.93-0.96 in both.

Notes

  • Made with SAGE, a mixed-precision quantization method for EXL3. The scores measure how closely the pack tracks the original checkpoint, not task accuracy, and no evaluation text was used to calibrate it.
  • The pack covers the text model: the 48-layer MoE backbone with its hybrid attention. MiMo's vision and audio encoders and its multi-token-prediction drafter are not part of the pack, so it serves text. The dflash drafter above is a separate repo.
  • The direction files are pinned by sha256 in the manifest and checked at load.
  • This build is intended for authorized red teaming and safety evaluation, the same scope as Blackfrost's release.