MiMo-V2.6-Flash-RL · DFlash draft model · EXL3 4.0 bpw
The DFlash draft model from XiaomiMiMo/MiMo-V2.6-Flash-RL,
converted to EXL3 at 4.0 bits per weight for speculative decoding under
exllamav3. It pairs with the
MiMo-V2.6-Flash-RL-EXL3 pack; the serving
recipe is MiMo-V2.6-Flash-RL-EXL3-recipe,
whose serve.sh loads it with -dm.
| Weights | 735 MB — 2.94 GB at BF16, so 4.0× smaller |
| Load with | -dm <this folder> (the runtime's draft-model flag) |
| Block size | 8, i.e. 7 drafted tokens per step |
| Taps | target layers 0 / 11 / 23 / 35 / 47, pinned as tap_shift: 0 in config.json |
| Fixes pre-applied | yes — see Packaging |
Why quantize a drafter
The drafter runs one forward per decode step, so its weights sit on the critical path: on a 524,288-token configuration it costs roughly 8 ms of a ~42 ms step. Cutting it to 735 MB measured +3.8% decode, and the 2.1 GB it frees raises the context ceiling too.
Measured on one RTX PRO 6000 with SixCat 0.7.0 speed against the 2.20 bpw pack, same
configuration in both runs (Q4 KV cache, 2048-token prefill chunks) with only the drafter swapped:
| Drafter | Size | Decode p50 | Context ceiling, Q4 KV | Context ceiling, FP16 KV |
|---|---|---|---|---|
| BF16 (corrected) | 2.94 GB | 193.66 tok/s | 524,288 | 139,264 |
| EXL3 4.0 bpw | 735 MB | 201.06 tok/s | 655,360 | 196,608 |
Draft acceptance is unchanged — the single-stream probe reports 0.25902668759811615 on every
run, both drafters.
It cannot change what the server outputs. The target verifies every drafted token, so under
greedy decoding the completion is identical. Verified rather than argued: same prompt, same
configuration, both drafters, 320 tokens — byte-identical output,
sha256 91f9c8811700a250d81ce1f3ed763252. The drafted tokens themselves do differ (acceptance
0.2925 for BF16 against 0.2804 here on that prompt), which is the expected behaviour: only speed
can move.
Files
| File | Size | What it is |
|---|---|---|
model.safetensors |
735 MB | The EXL3 weights, including the mask embedding as the tensor mask_embedding |
quantization_config.json |
49 KB | Bitrate, calibration mode and codebook per tensor |
config.json |
1.7 KB | Architecture DFlashDraftModel, block size, taps, tap_shift: 0 |
mask_embedding.safetensors |
8 KB | The learned mask embedding as a standalone one-tensor shard |
dflash.py |
14 KB | Reference implementation, referenced by config.json's auto_map |
| tokenizer files | 14 MB | Carried over so the folder can be re-converted standalone; serving uses the pack's |
Usage
hf download vcruz305/MiMo-V2.6-Flash-RL-dflash-EXL3-4.0bpw \
--local-dir MiMo-V2.6-Flash-RL-dflash-EXL3-4.0
# exllamav3, with the 2.20 bpw pack:
python examples/chat.py -m MiMo-V2.6-Flash-RL-EXL3/2.20bpw \
-dm MiMo-V2.6-Flash-RL-dflash-EXL3-4.0 -cs 65536
Or serve the measured configuration through the recipe:
PROFILE=q4kv-with-draft bash serve.sh # 655,360 context, 201.06 tok/s p50
Packaging: two details that matter
Both corrections that this architecture needs are already applied here, so no fix step is required before serving:
tap_shift: 0is pinned inconfig.json. The core port adds its own default of1totarget_layer_ids, which taps layers 1/12/24/36/48 instead of the correct 0/11/23/35/47 and collapses acceptance to ~0.04.- The mask embedding is a tensor inside the shard (key
mask_embedding) and also stands alone asmask_embedding.safetensors. The upstream checkpoint ships that vector only asmask_embedding.pt, which the runtime does not read, so mask positions fall back to an untrained embedding row and acceptance collapses the same way.
How it was built
convert.py -b 4.0 on the corrected BF16 drafter — about a minute on one GPU.
The drafter has no embedding table: fc.weight is [4096, 20480], a projection of the target's
five RMS-normed tapped hidden states, and the input layer embeds through the target's embedding.
The converter therefore cannot run its calibration forward pass and reports
Performing uncalibrated quantization, taking the side-model path it already uses for MTP heads
and vision towers. No calibration data is involved. tools/quantize_dflash.sh in the recipe
reproduces this build.
4.0 bpw was chosen over 8.0 bpw by measurement, not assumption: 8.0 bpw scored 67.03 tok/s against 68.77 for 4.0 on the same probe, at slightly lower acceptance — at 8 bpw the trellis unpack cost is not covered by the smaller byte count.