GLM-5.3-Flash DFlash2 Drafter — run dflash2G
Self-trained DFlash2 block-diffusion speculative-decoding drafter for GLM-5.3-Flash,
trained against our INT4 quant canada-quant/GLM-5.3-Flash-W4A16-MTP.
Trained on self-generated data only — no third-party drafter weights or traces anywhere in the training path. Supersedes
GLM-5.3-Flash-DFlash2-F and -E
(same architecture, same serve contract, drop-in). The first of our drafters to beat the reference
incoai/GLM-5.3-Flash-DFlash2 outside the protocol's noise on identical hardware:
3.676 vs 3.632 mean acceptance at K=7, 3.697 vs 3.667 single-stream (K=4: 3.088 vs 3.136, still 1.5% short). See below.
Architecture (unchanged from -E / -F)
| Field | Value |
|---|---|
| Type | DFlash2 block-diffusion drafter (DFlash2DraftModel, Qwen3-style backbone) |
| Layers | 8 decoder layers, full attention (no sliding window) |
| Hidden / heads | 4096 (intermediate 12288) · 32 attention heads / 8 KV heads · head_dim 128 |
| Target taps | 9 hidden-state taps at target layers [5, 9, 14, 19, 24, 28, 33, 38, 42] |
| Block size | 8 → K = 7 speculative tokens (num_speculative_tokens: 7) |
| Selector | rank 256, top_k 16, grouped dynamic conv (kernel 2, group 16) — trained, not vestigial |
| Mask embedding | learnable, shipped as mask_embedding.pt (mask_token_id 154856) |
| Size | ~1.84B drafter parameters · 6.2 GB bf16 checkpoint (ships untied embed_tokens + lm_head) |
| Max positions | 1,048,576 |
Training
- Warm-started from
dflash2F; 70,000 steps over 736,675 self-generated samples (the 449,600 of-F+ 291,020 new completions of never-before-generated prompts: real user chat, tool calling, hard math, code and STEM), 0.76 epoch, lr 1e-4 → 1e-5 (cosine), gamma-6.5 tail weighting, CE + top-20 KL, 8× NVIDIA B300, 27.4 h at 1.41 s/step. - Prompts sampled from public instruction sets — the
-Fset (ultrachat_200k · MIT, OpenR1-Math-220k · Apache-2.0, OpenMathReasoning · CC-BY-4.0, OpenScienceReasoning-2 · CC-BY-4.0, OpenCodeReasoning / OpenCodeInstruct · CC-BY-4.0, evol-codealpaca-v1 · Apache-2.0) plus, new for-G, lmsys-chat-1m and WildChat-1M (real user prompts, English, moderation-filtered), NuminaMath-1.5, OpenThoughts3 (math / code / science prompts), xlam-function-calling-60k (tool-calling prompts) and Nemotron-Post-Training-v2 STEM prompts; every completion was regenerated by the target model itself (thinking ON, T=1.0 / top_p 0.95, ≤ 4,096 tokens) — no third-party model outputs, only prompts. - The reference drafter incoai/GLM-5.3-Flash-DFlash2 was never a training input; it appears below only as a measured comparison.
PROVENANCE.txtis the verbatim training record. Recipe, code, patches and every raw measurement: canada-quant/vllm-glm53-flash-sm121/drafter (eval kit, every raw measurement, cards).
Holdout evaluation — same hardware, same day
500 never-trained-on prompts, thinking ON, T=1.0 / top_p 0.95, max_tokens 1024, greedy drafts, TP=4 (c16; the c1 row is 100 prompts at
concurrency 1). All rows below were measured on the same 8× B300 with the same vLLM build and the same target backend (FLASHINFER_MLA_SPARSE)
(-G on 2026-09-25; -F, -E and incoai on 2026-09-23) — acceptance length shifts by a few hundredths between GPU generations (incoai
reads 3.602–3.615 on H200 vs 3.632 on B300), so only same-hardware rows are compared.
| Mean acceptance length (output tok/s) | K=7, c16 | K=4, c16 | K=7, c1 |
|---|---|---|---|
dflash2G (this) |
3.676 (1,410) | 3.088 (1,331) | 3.697 (315) |
| incoai/GLM-5.3-Flash-DFlash2 (CC-BY-NC-ND-4.0) | 3.632 (1,425) | 3.136 (1,369) | 3.667 (328) |
dflash2F (our previous release) |
3.626 (1,406) | 3.085 (1,339) | 3.677 (313) |
dflash2E |
3.561 (1,397) | — | — |
Per-position acceptance (K=7, c16): 0.753 · 0.562 · 0.422 · 0.324 · 0.253 · 0.201 · 0.161 (incoai: 0.75 · 0.55 · 0.41 · 0.32 · 0.25 · 0.20 · 0.16;
-F: 0.750 · 0.555 · 0.415 · 0.316 · 0.244 · 0.193 · 0.153).
Deep traces (same protocol with max_tokens 4096, K=7, c16): this drafter 3.478, incoai 3.473 — parity. Both drafters lose
~0.16–0.20 acceptance when the thinking trace runs to 4K tokens, and the 1K-token lead above does not carry that deep; quote both numbers.
(Per-position @4K: 0.729 · 0.526 · 0.388 · 0.292 · 0.225 · 0.177 · 0.141 vs incoai 0.728 · 0.523 · 0.384 · 0.291 · 0.226 · 0.179 · 0.144.)
Honest read: +0.044 over the reference at K=7 and +0.030 single-stream, both outside the protocol's ±0.015 run-to-run noise — every
draft position is at or above the reference. K=4 is still 1.5% short (−0.048): the reference keeps a small edge on the first two positions
when only four are drafted. Against -F it is +0.050 at K=7 from 291K fresh, more diverse self-generated samples (real chat and tool-calling
prompts the earlier sets never had). Fully ours (Apache-2.0, no NC/ND terms), trained on data we control, reproducible end-to-end.
DGX Spark (SM121): -E reproduced its H200 acceptance on 2× Spark within 0.3% (3.5788 c16 / 3.5970 c1 vs 3.568 / 3.585). -G is a
drop-in for the same launcher and the Spark drafter of record since 2026-09-25: the H2H champion leg (2026-09-27; 702/702 requests, coherence-clean; tg32 d0 31.35±3.52 / c1@65k 15.34 quiescent — receipt: H2H_FINAL_VERDICT_2026-09-27.md, spark-cluster PR #27) and the live production serve both ran G (sha256-verified on both hosts). A same-rig Spark acceptance A/B vs the reference remains unmeasured (the acceptance table above is B300) — re-measure
with the eval kit in the repo before quoting a Spark acceptance number.
Serving
Pairs with canada-quant/GLM-5.3-Flash-W4A16-MTP (the quant it was trained against) or BF16 GLM-5.3-Flash. vLLM speculative config:
{"method": "dflash", "model": "/models/GLM-5.3-Flash-DFlash2-G", "num_speculative_tokens": 7}
x86 (H100 / H200 / B300), upstream vLLM nightly ≥ 2026-09-08 (+ the two GLM-5.3-Flash DFlash2 bind-mount patches in the repo above):
vllm serve /models/glm53-flash-w4a16-mtp --served-model-name glm53 --tensor-parallel-size 4 --enable-expert-parallel \
--block-size 64 --no-enable-prefix-caching \
--speculative-config '{"method":"dflash","model":"/models/GLM-5.3-Flash-DFlash2-G","num_speculative_tokens":7}'
2× DGX Spark (SM121): prebuilt image + one-command launcher in
canada-quant/vllm-glm53-flash-sm121 — the drafter is bind-mounted at runtime
(DRAFTER_HOST_PATH=/models/GLM-5.3-Flash-DFlash2-G), no rebuild. launch_dflash2_tp2.sh in this repo is the
same TP=2 launcher with this drafter as its default.
Hard constraints (all measured, not stylistic):
num_speculative_tokensmust be 7 (= block_size − 1). Other counts boot-wedge the DFlash2 stack.mask_embedding.ptmust sit next to the weights. Verify the boot log carriesLoaded DFlash mask embedding for mask_token_id 154856 from mask_embedding.pt— absence means the mask was silently ignored; do not serve.- The 9-tap config requires the serving stack to honor
dflash_config.target_layer_idsof length 9 (upstream vLLM DFlash2 does — vllm-project/vllm#52816). - Full-attention drafter layers: the target stack must accept
FullAttentionSpecdrafter KV in the GLM-5 KV fast path (the SM121 image above has it; upstream nightly needs thekv_cache_utilspatch from the repo).
Provenance
- Files:
model.safetensors(sha2560d2a3eebfc532adfe9d092f0b972780592be8b587cf0fcc871d6c936c0b353f6),config.json(699bdf29…c51778, identical to-E/-F),mask_embedding.pt(291c7243…ec82ec),PROVENANCE.txt. - Every number on this card comes from sha256-pinned raw eval JSONs banked in the repo (
angelspec-glm53/runs/all/dflash2G-*.jsonincl.dflash2G-k7-mt4096.json,incoai-k7-mt4096-b300.json,dflash2F-*.json,dflash2E-k7-b300.json,incoai-k7-b300.json,incoai-k4-b300.json,incoai-k7-c1-b300.json).
References
- DFlash2 method: z-lab/dflash · arXiv 2602.06036 · inco.ai blog
- Upstream vLLM DFlash2 integration: vllm-project/vllm#52816
- Trainer: Tencent/AngelSpec (Apache-2.0)
- Target model: zai-org/GLM-5.3-Flash · our quant: canada-quant/GLM-5.3-Flash-W4A16-MTP
- Reference drafter (comparison only): incoai/GLM-5.3-Flash-DFlash2 (CC-BY-NC-ND-4.0)