canada-quant/GLM-5.3-Flash-DFlash2-G

🤗 Hugging Face 来源text-generationapache-2.03.1B 参数6.2 GBsafetensors✓ 2 个校验和今天更新
需要做种者 →

GLM-5.3-Flash DFlash2 Drafter — run dflash2G

Self-trained DFlash2 block-diffusion speculative-decoding drafter for GLM-5.3-Flash, trained against our INT4 quant canada-quant/GLM-5.3-Flash-W4A16-MTP. Trained on self-generated data only — no third-party drafter weights or traces anywhere in the training path. Supersedes GLM-5.3-Flash-DFlash2-F and -E (same architecture, same serve contract, drop-in). The first of our drafters to beat the reference incoai/GLM-5.3-Flash-DFlash2 outside the protocol's noise on identical hardware: 3.676 vs 3.632 mean acceptance at K=7, 3.697 vs 3.667 single-stream (K=4: 3.088 vs 3.136, still 1.5% short). See below.

Architecture (unchanged from -E / -F)

Field Value
Type DFlash2 block-diffusion drafter (DFlash2DraftModel, Qwen3-style backbone)
Layers 8 decoder layers, full attention (no sliding window)
Hidden / heads 4096 (intermediate 12288) · 32 attention heads / 8 KV heads · head_dim 128
Target taps 9 hidden-state taps at target layers [5, 9, 14, 19, 24, 28, 33, 38, 42]
Block size 8 → K = 7 speculative tokens (num_speculative_tokens: 7)
Selector rank 256, top_k 16, grouped dynamic conv (kernel 2, group 16) — trained, not vestigial
Mask embedding learnable, shipped as mask_embedding.pt (mask_token_id 154856)
Size ~1.84B drafter parameters · 6.2 GB bf16 checkpoint (ships untied embed_tokens + lm_head)
Max positions 1,048,576

Training

  • Warm-started from dflash2F; 70,000 steps over 736,675 self-generated samples (the 449,600 of -F + 291,020 new completions of never-before-generated prompts: real user chat, tool calling, hard math, code and STEM), 0.76 epoch, lr 1e-4 → 1e-5 (cosine), gamma-6.5 tail weighting, CE + top-20 KL, 8× NVIDIA B300, 27.4 h at 1.41 s/step.
  • Prompts sampled from public instruction sets — the -F set (ultrachat_200k · MIT, OpenR1-Math-220k · Apache-2.0, OpenMathReasoning · CC-BY-4.0, OpenScienceReasoning-2 · CC-BY-4.0, OpenCodeReasoning / OpenCodeInstruct · CC-BY-4.0, evol-codealpaca-v1 · Apache-2.0) plus, new for -G, lmsys-chat-1m and WildChat-1M (real user prompts, English, moderation-filtered), NuminaMath-1.5, OpenThoughts3 (math / code / science prompts), xlam-function-calling-60k (tool-calling prompts) and Nemotron-Post-Training-v2 STEM prompts; every completion was regenerated by the target model itself (thinking ON, T=1.0 / top_p 0.95, ≤ 4,096 tokens) — no third-party model outputs, only prompts.
  • The reference drafter incoai/GLM-5.3-Flash-DFlash2 was never a training input; it appears below only as a measured comparison.
  • PROVENANCE.txt is the verbatim training record. Recipe, code, patches and every raw measurement: canada-quant/vllm-glm53-flash-sm121/drafter (eval kit, every raw measurement, cards).

Holdout evaluation — same hardware, same day

500 never-trained-on prompts, thinking ON, T=1.0 / top_p 0.95, max_tokens 1024, greedy drafts, TP=4 (c16; the c1 row is 100 prompts at concurrency 1). All rows below were measured on the same 8× B300 with the same vLLM build and the same target backend (FLASHINFER_MLA_SPARSE) (-G on 2026-09-25; -F, -E and incoai on 2026-09-23) — acceptance length shifts by a few hundredths between GPU generations (incoai reads 3.602–3.615 on H200 vs 3.632 on B300), so only same-hardware rows are compared.

Mean acceptance length (output tok/s) K=7, c16 K=4, c16 K=7, c1
dflash2G (this) 3.676 (1,410) 3.088 (1,331) 3.697 (315)
incoai/GLM-5.3-Flash-DFlash2 (CC-BY-NC-ND-4.0) 3.632 (1,425) 3.136 (1,369) 3.667 (328)
dflash2F (our previous release) 3.626 (1,406) 3.085 (1,339) 3.677 (313)
dflash2E 3.561 (1,397) — —

Per-position acceptance (K=7, c16): 0.753 · 0.562 · 0.422 · 0.324 · 0.253 · 0.201 · 0.161 (incoai: 0.75 · 0.55 · 0.41 · 0.32 · 0.25 · 0.20 · 0.16; -F: 0.750 · 0.555 · 0.415 · 0.316 · 0.244 · 0.193 · 0.153).

Deep traces (same protocol with max_tokens 4096, K=7, c16): this drafter 3.478, incoai 3.473 — parity. Both drafters lose ~0.16–0.20 acceptance when the thinking trace runs to 4K tokens, and the 1K-token lead above does not carry that deep; quote both numbers. (Per-position @4K: 0.729 · 0.526 · 0.388 · 0.292 · 0.225 · 0.177 · 0.141 vs incoai 0.728 · 0.523 · 0.384 · 0.291 · 0.226 · 0.179 · 0.144.)

Honest read: +0.044 over the reference at K=7 and +0.030 single-stream, both outside the protocol's ±0.015 run-to-run noise — every draft position is at or above the reference. K=4 is still 1.5% short (−0.048): the reference keeps a small edge on the first two positions when only four are drafted. Against -F it is +0.050 at K=7 from 291K fresh, more diverse self-generated samples (real chat and tool-calling prompts the earlier sets never had). Fully ours (Apache-2.0, no NC/ND terms), trained on data we control, reproducible end-to-end.

DGX Spark (SM121): -E reproduced its H200 acceptance on 2× Spark within 0.3% (3.5788 c16 / 3.5970 c1 vs 3.568 / 3.585). -G is a drop-in for the same launcher and the Spark drafter of record since 2026-09-25: the H2H champion leg (2026-09-27; 702/702 requests, coherence-clean; tg32 d0 31.35±3.52 / c1@65k 15.34 quiescent — receipt: H2H_FINAL_VERDICT_2026-09-27.md, spark-cluster PR #27) and the live production serve both ran G (sha256-verified on both hosts). A same-rig Spark acceptance A/B vs the reference remains unmeasured (the acceptance table above is B300) — re-measure with the eval kit in the repo before quoting a Spark acceptance number.

Serving

Pairs with canada-quant/GLM-5.3-Flash-W4A16-MTP (the quant it was trained against) or BF16 GLM-5.3-Flash. vLLM speculative config:

{"method": "dflash", "model": "/models/GLM-5.3-Flash-DFlash2-G", "num_speculative_tokens": 7}

x86 (H100 / H200 / B300), upstream vLLM nightly ≥ 2026-09-08 (+ the two GLM-5.3-Flash DFlash2 bind-mount patches in the repo above):

vllm serve /models/glm53-flash-w4a16-mtp --served-model-name glm53 --tensor-parallel-size 4 --enable-expert-parallel \
  --block-size 64 --no-enable-prefix-caching \
  --speculative-config '{"method":"dflash","model":"/models/GLM-5.3-Flash-DFlash2-G","num_speculative_tokens":7}'

2× DGX Spark (SM121): prebuilt image + one-command launcher in canada-quant/vllm-glm53-flash-sm121 — the drafter is bind-mounted at runtime (DRAFTER_HOST_PATH=/models/GLM-5.3-Flash-DFlash2-G), no rebuild. launch_dflash2_tp2.sh in this repo is the same TP=2 launcher with this drafter as its default.

Hard constraints (all measured, not stylistic):

  • num_speculative_tokens must be 7 (= block_size − 1). Other counts boot-wedge the DFlash2 stack.
  • mask_embedding.pt must sit next to the weights. Verify the boot log carries Loaded DFlash mask embedding for mask_token_id 154856 from mask_embedding.pt — absence means the mask was silently ignored; do not serve.
  • The 9-tap config requires the serving stack to honor dflash_config.target_layer_ids of length 9 (upstream vLLM DFlash2 does — vllm-project/vllm#52816).
  • Full-attention drafter layers: the target stack must accept FullAttentionSpec drafter KV in the GLM-5 KV fast path (the SM121 image above has it; upstream nightly needs the kv_cache_utils patch from the repo).

Provenance

  • Files: model.safetensors (sha256 0d2a3eebfc532adfe9d092f0b972780592be8b587cf0fcc871d6c936c0b353f6), config.json (699bdf29…c51778, identical to -E / -F), mask_embedding.pt (291c7243…ec82ec), PROVENANCE.txt.
  • Every number on this card comes from sha256-pinned raw eval JSONs banked in the repo (angelspec-glm53/runs/all/dflash2G-*.json incl. dflash2G-k7-mt4096.json, incoai-k7-mt4096-b300.json, dflash2F-*.json, dflash2E-k7-b300.json, incoai-k7-b300.json, incoai-k4-b300.json, incoai-k7-c1-b300.json).

References