canada-quant/GLM-5.3-Flash-DFlash2-F

🤗 Hugging Face 来源text-generationapache-2.03.1B 参数6.2 GBsafetensors✓ 2 个校验和今天更新
需要做种者 →

GLM-5.3-Flash DFlash2 Drafter — run dflash2F

Self-trained DFlash2 block-diffusion speculative-decoding drafter for GLM-5.3-Flash, trained against our INT4 quant canada-quant/GLM-5.3-Flash-W4A16-MTP. Trained on self-generated data only — no third-party drafter weights or traces anywhere in the training path. Supersedes GLM-5.3-Flash-DFlash2-E (same architecture, same serve contract, drop-in): +0.065 mean acceptance at K=7 on the same hardware, and parity with the reference drafter incoai/GLM-5.3-Flash-DFlash2 measured on the same GPUs the same day (see below).

Architecture (unchanged from -E)

Field Value
Type DFlash2 block-diffusion drafter (DFlash2DraftModel, Qwen3-style backbone)
Layers 8 decoder layers, full attention (no sliding window)
Hidden / heads 4096 (intermediate 12288) · 32 attention heads / 8 KV heads · head_dim 128
Target taps 9 hidden-state taps at target layers [5, 9, 14, 19, 24, 28, 33, 38, 42]
Block size 8 → K = 7 speculative tokens (num_speculative_tokens: 7)
Selector rank 256, top_k 16, grouped dynamic conv (kernel 2, group 16) — trained, not vestigial
Mask embedding learnable, shipped as mask_embedding.pt (mask_token_id 154856)
Size ~1.84B drafter parameters · 6.2 GB bf16 checkpoint (ships untied embed_tokens + lm_head)
Max positions 1,048,576

Training

  • Warm-started from dflash2E; one pass over 449,600 self-generated samples (the 350,260 of -E + 99,340 new completions of never-before-generated prompts), 41,279 steps, lr 1e-4 → 1e-5 (cosine), gamma-6.5 tail weighting, CE + top-20 KL, 8× NVIDIA B300, 16.5 h.
  • Prompts sampled from public instruction sets (ultrachat_200k · MIT, OpenR1-Math-220k · Apache-2.0, OpenMathReasoning · CC-BY-4.0, OpenScienceReasoning-2 · CC-BY-4.0, OpenCodeReasoning / OpenCodeInstruct · CC-BY-4.0, evol-codealpaca-v1 · Apache-2.0); every completion was regenerated by the target model itself (thinking ON, T=1.0 / top_p 0.95, ≤ 4,096 tokens) — no third-party model outputs.
  • The reference drafter incoai/GLM-5.3-Flash-DFlash2 was never a training input; it appears below only as a measured comparison.
  • PROVENANCE.txt is the verbatim training record. Recipe, code, patches and every raw measurement: canada-quant/vllm-glm53-flash-sm121/drafter (eval kit, every raw measurement, cards).

Holdout evaluation — same hardware, same day

500 never-trained-on prompts, thinking ON, T=1.0 / top_p 0.95, max_tokens 1024, greedy drafts, TP=4 (c16; the c1 row is 100 prompts at concurrency 1). All rows below were measured on the same 8× B300 within six hours of each other, with the same vLLM build and the same target backend (FLASHINFER_MLA_SPARSE) — acceptance length shifts by a few hundredths between GPU generations (incoai reads 3.602–3.615 on H200 vs 3.632 on B300), so only same-hardware rows are compared.

Mean acceptance length (output tok/s) K=7, c16 K=4, c16 K=7, c1
dflash2F (this) 3.626 (1,406) 3.085 (1,339) 3.677 (313)
incoai/GLM-5.3-Flash-DFlash2 (CC-BY-NC-ND-4.0) 3.632 (1,425) 3.136 (1,369) 3.667 (328)
dflash2E (our previous release) 3.561 (1,397) — —

Per-position acceptance (K=7, c16): 0.750 · 0.555 · 0.415 · 0.316 · 0.244 · 0.193 · 0.153 (incoai: 0.75 · 0.55 · 0.41 · 0.32 · 0.25 · 0.20 · 0.16).

Honest read: parity with the reference at K=7 (−0.006) and single-stream (+0.010), both inside the protocol's ±0.015 run-to-run noise; 1.6% short at K=4 (−0.051, the first draft positions). Against -E it is +0.065 at K=7 — the largest single-leg gain in the lineage, from 100K fresh self-generated samples. We ship it because it is fully ours (Apache-2.0, no NC/ND terms), trained on data we control, reproducible end-to-end, and now level with the reference. On the H200 (previous campaign) -E measured 3.568 / 3.067 / 3.585.

DGX Spark (SM121): -E reproduced its H200 acceptance on 2× Spark within 0.3% (3.5788 c16 / 3.5970 c1 vs 3.568 / 3.585). -F is a drop-in for the same launcher; its Spark same-rig A/B against -E and the reference is the next measurement — re-measure both drafters on your rig with the eval kit in the repo before quoting a Spark number.

Serving

Pairs with canada-quant/GLM-5.3-Flash-W4A16-MTP (the quant it was trained against) or BF16 GLM-5.3-Flash. vLLM speculative config:

{"method": "dflash", "model": "/models/GLM-5.3-Flash-DFlash2-F", "num_speculative_tokens": 7}

x86 (H100 / H200 / B300), upstream vLLM nightly ≥ 2026-09-08 (+ the two GLM-5.3-Flash DFlash2 bind-mount patches in the repo above):

vllm serve /models/glm53-flash-w4a16-mtp --served-model-name glm53 --tensor-parallel-size 4 --enable-expert-parallel \
  --block-size 64 --no-enable-prefix-caching \
  --speculative-config '{"method":"dflash","model":"/models/GLM-5.3-Flash-DFlash2-F","num_speculative_tokens":7}'

2× DGX Spark (SM121): prebuilt image + one-command launcher in canada-quant/vllm-glm53-flash-sm121 — the drafter is bind-mounted at runtime (DRAFTER_HOST_PATH=/models/GLM-5.3-Flash-DFlash2-F), no rebuild. launch_dflash2_tp2.sh in this repo is the same TP=2 launcher with this drafter as its default.

Hard constraints (all measured, not stylistic):

  • num_speculative_tokens must be 7 (= block_size − 1). Other counts boot-wedge the DFlash2 stack.
  • mask_embedding.pt must sit next to the weights. Verify the boot log carries Loaded DFlash mask embedding for mask_token_id 154856 from mask_embedding.pt — absence means the mask was silently ignored; do not serve.
  • The 9-tap config requires the serving stack to honor dflash_config.target_layer_ids of length 9 (upstream vLLM DFlash2 does — vllm-project/vllm#52816).
  • Full-attention drafter layers: the target stack must accept FullAttentionSpec drafter KV in the GLM-5 KV fast path (the SM121 image above has it; upstream nightly needs the kv_cache_utils patch from the repo).

Provenance

  • Files: model.safetensors (sha256 305efe2aa50447cbac7d7d3ab566985103973a79c5d21eb11e1dbf870139dd8c), config.json (699bdf29…c51778, identical to -E), mask_embedding.pt (0a5eecd9…304f3a), PROVENANCE.txt (546442a6…f75721).
  • Every number on this card comes from sha256-pinned raw eval JSONs banked in the repo (angelspec-glm53/runs/all/dflash2F-*.json, dflash2E-k7-b300.json, incoai-k7-b300.json, incoai-k4-b300.json, incoai-k7-c1-b300.json).

References