GLM-5.3-Flash DFlash2 Drafter — run dflash2F
Self-trained DFlash2 block-diffusion speculative-decoding drafter for GLM-5.3-Flash,
trained against our INT4 quant canada-quant/GLM-5.3-Flash-W4A16-MTP.
Trained on self-generated data only — no third-party drafter weights or traces anywhere in the training path. Supersedes
GLM-5.3-Flash-DFlash2-E (same architecture, same serve contract, drop-in):
+0.065 mean acceptance at K=7 on the same hardware, and parity with the reference drafter
incoai/GLM-5.3-Flash-DFlash2 measured on the same GPUs the same day (see below).
Architecture (unchanged from -E)
| Field | Value |
|---|---|
| Type | DFlash2 block-diffusion drafter (DFlash2DraftModel, Qwen3-style backbone) |
| Layers | 8 decoder layers, full attention (no sliding window) |
| Hidden / heads | 4096 (intermediate 12288) · 32 attention heads / 8 KV heads · head_dim 128 |
| Target taps | 9 hidden-state taps at target layers [5, 9, 14, 19, 24, 28, 33, 38, 42] |
| Block size | 8 → K = 7 speculative tokens (num_speculative_tokens: 7) |
| Selector | rank 256, top_k 16, grouped dynamic conv (kernel 2, group 16) — trained, not vestigial |
| Mask embedding | learnable, shipped as mask_embedding.pt (mask_token_id 154856) |
| Size | ~1.84B drafter parameters · 6.2 GB bf16 checkpoint (ships untied embed_tokens + lm_head) |
| Max positions | 1,048,576 |
Training
- Warm-started from
dflash2E; one pass over 449,600 self-generated samples (the 350,260 of-E+ 99,340 new completions of never-before-generated prompts), 41,279 steps, lr 1e-4 → 1e-5 (cosine), gamma-6.5 tail weighting, CE + top-20 KL, 8× NVIDIA B300, 16.5 h. - Prompts sampled from public instruction sets (ultrachat_200k · MIT, OpenR1-Math-220k · Apache-2.0, OpenMathReasoning · CC-BY-4.0, OpenScienceReasoning-2 · CC-BY-4.0, OpenCodeReasoning / OpenCodeInstruct · CC-BY-4.0, evol-codealpaca-v1 · Apache-2.0); every completion was regenerated by the target model itself (thinking ON, T=1.0 / top_p 0.95, ≤ 4,096 tokens) — no third-party model outputs.
- The reference drafter incoai/GLM-5.3-Flash-DFlash2 was never a training input; it appears below only as a measured comparison.
PROVENANCE.txtis the verbatim training record. Recipe, code, patches and every raw measurement: canada-quant/vllm-glm53-flash-sm121/drafter (eval kit, every raw measurement, cards).
Holdout evaluation — same hardware, same day
500 never-trained-on prompts, thinking ON, T=1.0 / top_p 0.95, max_tokens 1024, greedy drafts, TP=4 (c16; the c1 row is 100 prompts at
concurrency 1). All rows below were measured on the same 8× B300 within six hours of each other, with the same vLLM build and the same
target backend (FLASHINFER_MLA_SPARSE) — acceptance length shifts by a few hundredths between GPU generations (incoai reads 3.602–3.615
on H200 vs 3.632 on B300), so only same-hardware rows are compared.
| Mean acceptance length (output tok/s) | K=7, c16 | K=4, c16 | K=7, c1 |
|---|---|---|---|
dflash2F (this) |
3.626 (1,406) | 3.085 (1,339) | 3.677 (313) |
| incoai/GLM-5.3-Flash-DFlash2 (CC-BY-NC-ND-4.0) | 3.632 (1,425) | 3.136 (1,369) | 3.667 (328) |
dflash2E (our previous release) |
3.561 (1,397) | — | — |
Per-position acceptance (K=7, c16): 0.750 · 0.555 · 0.415 · 0.316 · 0.244 · 0.193 · 0.153 (incoai: 0.75 · 0.55 · 0.41 · 0.32 · 0.25 · 0.20 · 0.16).
Honest read: parity with the reference at K=7 (−0.006) and single-stream (+0.010), both inside the protocol's ±0.015 run-to-run noise;
1.6% short at K=4 (−0.051, the first draft positions). Against -E it is +0.065 at K=7 — the largest single-leg gain in the lineage, from
100K fresh self-generated samples. We ship it because it is fully ours (Apache-2.0, no NC/ND terms), trained on data we control,
reproducible end-to-end, and now level with the reference. On the H200 (previous campaign) -E measured 3.568 / 3.067 / 3.585.
DGX Spark (SM121): -E reproduced its H200 acceptance on 2× Spark within 0.3% (3.5788 c16 / 3.5970 c1 vs 3.568 / 3.585). -F is a
drop-in for the same launcher; its Spark same-rig A/B against -E and the reference is the next measurement — re-measure both drafters on
your rig with the eval kit in the repo before quoting a Spark number.
Serving
Pairs with canada-quant/GLM-5.3-Flash-W4A16-MTP (the quant it was trained against) or BF16 GLM-5.3-Flash. vLLM speculative config:
{"method": "dflash", "model": "/models/GLM-5.3-Flash-DFlash2-F", "num_speculative_tokens": 7}
x86 (H100 / H200 / B300), upstream vLLM nightly ≥ 2026-09-08 (+ the two GLM-5.3-Flash DFlash2 bind-mount patches in the repo above):
vllm serve /models/glm53-flash-w4a16-mtp --served-model-name glm53 --tensor-parallel-size 4 --enable-expert-parallel \
--block-size 64 --no-enable-prefix-caching \
--speculative-config '{"method":"dflash","model":"/models/GLM-5.3-Flash-DFlash2-F","num_speculative_tokens":7}'
2× DGX Spark (SM121): prebuilt image + one-command launcher in
canada-quant/vllm-glm53-flash-sm121 — the drafter is bind-mounted at runtime
(DRAFTER_HOST_PATH=/models/GLM-5.3-Flash-DFlash2-F), no rebuild. launch_dflash2_tp2.sh in this repo is the
same TP=2 launcher with this drafter as its default.
Hard constraints (all measured, not stylistic):
num_speculative_tokensmust be 7 (= block_size − 1). Other counts boot-wedge the DFlash2 stack.mask_embedding.ptmust sit next to the weights. Verify the boot log carriesLoaded DFlash mask embedding for mask_token_id 154856 from mask_embedding.pt— absence means the mask was silently ignored; do not serve.- The 9-tap config requires the serving stack to honor
dflash_config.target_layer_idsof length 9 (upstream vLLM DFlash2 does — vllm-project/vllm#52816). - Full-attention drafter layers: the target stack must accept
FullAttentionSpecdrafter KV in the GLM-5 KV fast path (the SM121 image above has it; upstream nightly needs thekv_cache_utilspatch from the repo).
Provenance
- Files:
model.safetensors(sha256305efe2aa50447cbac7d7d3ab566985103973a79c5d21eb11e1dbf870139dd8c),config.json(699bdf29…c51778, identical to-E),mask_embedding.pt(0a5eecd9…304f3a),PROVENANCE.txt(546442a6…f75721). - Every number on this card comes from sha256-pinned raw eval JSONs banked in the repo (
angelspec-glm53/runs/all/dflash2F-*.json,dflash2E-k7-b300.json,incoai-k7-b300.json,incoai-k4-b300.json,incoai-k7-c1-b300.json).
References
- DFlash2 method: z-lab/dflash · arXiv 2602.06036 · inco.ai blog
- Upstream vLLM DFlash2 integration: vllm-project/vllm#52816
- Trainer: Tencent/AngelSpec (Apache-2.0)
- Target model: zai-org/GLM-5.3-Flash · our quant: canada-quant/GLM-5.3-Flash-W4A16-MTP
- Reference drafter (comparison only): incoai/GLM-5.3-Flash-DFlash2 (CC-BY-NC-ND-4.0)