vcruz305/MiMo-V2.6-Flash-RL-dflash-EXL3-4.0bpw

🤗 Hugging Face sourcetext-generationmit368M params735 MBsafetensors✓ 3 checksumsupdated today
Needs seeder →

MiMo-V2.6-Flash-RL · DFlash draft model · EXL3 4.0 bpw

The DFlash draft model from XiaomiMiMo/MiMo-V2.6-Flash-RL, converted to EXL3 at 4.0 bits per weight for speculative decoding under exllamav3. It pairs with the MiMo-V2.6-Flash-RL-EXL3 pack; the serving recipe is MiMo-V2.6-Flash-RL-EXL3-recipe, whose serve.sh loads it with -dm.

Weights 735 MB — 2.94 GB at BF16, so 4.0× smaller
Load with -dm <this folder> (the runtime's draft-model flag)
Block size 8, i.e. 7 drafted tokens per step
Taps target layers 0 / 11 / 23 / 35 / 47, pinned as tap_shift: 0 in config.json
Fixes pre-applied yes — see Packaging

Why quantize a drafter

The drafter runs one forward per decode step, so its weights sit on the critical path: on a 524,288-token configuration it costs roughly 8 ms of a ~42 ms step. Cutting it to 735 MB measured +3.8% decode, and the 2.1 GB it frees raises the context ceiling too.

Measured on one RTX PRO 6000 with SixCat 0.7.0 speed against the 2.20 bpw pack, same configuration in both runs (Q4 KV cache, 2048-token prefill chunks) with only the drafter swapped:

Drafter Size Decode p50 Context ceiling, Q4 KV Context ceiling, FP16 KV
BF16 (corrected) 2.94 GB 193.66 tok/s 524,288 139,264
EXL3 4.0 bpw 735 MB 201.06 tok/s 655,360 196,608

Draft acceptance is unchanged — the single-stream probe reports 0.25902668759811615 on every run, both drafters.

It cannot change what the server outputs. The target verifies every drafted token, so under greedy decoding the completion is identical. Verified rather than argued: same prompt, same configuration, both drafters, 320 tokens — byte-identical output, sha256 91f9c8811700a250d81ce1f3ed763252. The drafted tokens themselves do differ (acceptance 0.2925 for BF16 against 0.2804 here on that prompt), which is the expected behaviour: only speed can move.

Files

File Size What it is
model.safetensors 735 MB The EXL3 weights, including the mask embedding as the tensor mask_embedding
quantization_config.json 49 KB Bitrate, calibration mode and codebook per tensor
config.json 1.7 KB Architecture DFlashDraftModel, block size, taps, tap_shift: 0
mask_embedding.safetensors 8 KB The learned mask embedding as a standalone one-tensor shard
dflash.py 14 KB Reference implementation, referenced by config.json's auto_map
tokenizer files 14 MB Carried over so the folder can be re-converted standalone; serving uses the pack's

Usage

hf download vcruz305/MiMo-V2.6-Flash-RL-dflash-EXL3-4.0bpw \
    --local-dir MiMo-V2.6-Flash-RL-dflash-EXL3-4.0

# exllamav3, with the 2.20 bpw pack:
python examples/chat.py -m MiMo-V2.6-Flash-RL-EXL3/2.20bpw \
    -dm MiMo-V2.6-Flash-RL-dflash-EXL3-4.0 -cs 65536

Or serve the measured configuration through the recipe:

PROFILE=q4kv-with-draft bash serve.sh     # 655,360 context, 201.06 tok/s p50

Packaging: two details that matter

Both corrections that this architecture needs are already applied here, so no fix step is required before serving:

  1. tap_shift: 0 is pinned in config.json. The core port adds its own default of 1 to target_layer_ids, which taps layers 1/12/24/36/48 instead of the correct 0/11/23/35/47 and collapses acceptance to ~0.04.
  2. The mask embedding is a tensor inside the shard (key mask_embedding) and also stands alone as mask_embedding.safetensors. The upstream checkpoint ships that vector only as mask_embedding.pt, which the runtime does not read, so mask positions fall back to an untrained embedding row and acceptance collapses the same way.

How it was built

convert.py -b 4.0 on the corrected BF16 drafter — about a minute on one GPU.

The drafter has no embedding table: fc.weight is [4096, 20480], a projection of the target's five RMS-normed tapped hidden states, and the input layer embeds through the target's embedding. The converter therefore cannot run its calibration forward pass and reports Performing uncalibrated quantization, taking the side-model path it already uses for MTP heads and vision towers. No calibration data is involved. tools/quantize_dflash.sh in the recipe reproduces this build.

4.0 bpw was chosen over 8.0 bpw by measurement, not assumption: 8.0 bpw scored 67.03 tok/s against 68.77 for 4.0 on the same probe, at slightly lower acceptance — at 8 bpw the trellis unpack cost is not covered by the smaller byte count.