canada-quant/GLM-5.3-Flash-W4A16-MTP

🤗 Hugging Face 来源image-text-to-textmit49.9B 参数1.3 TBsafetensors✓ 13 个校验和今天更新
需要做种者 →

GLM-5.3-Flash — W4A16 (INT4) + BF16 MTP

Model description

INT4 weight-only quantization of zai-org/GLM-5.3-Flash with the BF16 MTP draft head kept for speculative decoding. Only the 36,288 routed-expert GEMMs are INT4 (GPTQ, symmetric, group-size 128); attention, router, shared experts, embeddings, the vision tower and the MTP head stay in BF16. Not a new model — all capability comes from the base model.

Size 177.7 GiB (BF16 ≈ 599 GiB, −70%)
Runs on NVIDIA H100, H200, RTX PRO 6000, and DGX Spark (GB10) — validated GPU counts and context per config in Hardware & context limits
Context full 1,048,576 tokens on 2× DGX Spark and 2× H200; 262K–512K on the 4/8-GPU x86 configs (KV-memory-bound)
Quality AIME 2025 0.8833 (n=120) vs 0.9000 for the NVFP4 reference on H100, within noise; GSM8K 0.97; GPQA-Diamond within noise; vision: MMMU 0.747 vs 0.68 NVFP4 (same rig) · OCRBench 888 vs 882
Throughput 8× H100: 181.2 tok/s single-stream latency recipe (TP8 + DFlash2-G K=7, +46%), 2,418 tok/s @c256 aggregate (TP8 spec-off); 4× H200 TP=4: 195 → 1,954 tok/s from c1 to c256; beats NVIDIA's NVFP4 checkpoint on H100 (+47.8% c1, +5.1% c256); matches the NVFP4 reference on RTX PRO 6000

Full benchmark grids, comparison protocols and research notes: BENCHMARKS.md.

Uses & recommended recipes

Quick start

# 1. Download (~178 GiB)
huggingface-cli download canada-quant/GLM-5.3-Flash-W4A16-MTP --local-dir /models/glm53-flash-w4a16-mtp

# 2. Serve on 4× H100 / H200 (other hardware: see Serving recipes)
docker run --gpus '"device=0,1,2,3"' --ipc=host --network=host --rm \
  -v /models:/models vllm/vllm-openai:glm53-flash-x86_64-cu130 \
  vllm serve /models/glm53-flash-w4a16-mtp --served-model-name glm53-w4 \
    --tensor-parallel-size 4 --enable-expert-parallel \
    --max-model-len 262144 --max-num-seqs 512 \
    --gpu-memory-utilization 0.92 --no-enable-prefix-caching \
    --speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
    --reasoning-parser glm45 --tool-call-parser glm47 --enable-auto-tool-choice \
    --trust-remote-code --port 8000

# 3. Call it (OpenAI-compatible)
curl http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "glm53-w4",
  "messages": [{"role": "user", "content": "Prove there are infinitely many primes."}]
}'

Two things that bite. If your config.json predates 2026-09-08, re-download it — older copies fail in vLLM with KeyError: 'layers.0.mlp.gate_up_proj.weight' (weights are unchanged). And always pass --max-num-seqs ≤ 512 — the vLLM default of 1024 exceeds this hybrid linear-attention model's 512 Mamba/KDA-state cache blocks.

Serving recipes

Architecture Image
SM90 (H100 / H200) vllm/vllm-openai:glm53-flash-x86_64-cu130 (validated). Upstream vllm/vllm-openai:nightly-x86_64 ≥ 2026-09-08 also boots this checkpoint (vllm-project/vllm#53906); pass --attention-backend FLASH_ATTN_MLA_SPARSE there, its default backend faults on ≥131K prompts.
SM90, 8× H100 + DFlash2-G drafter ghcr.io/canada-quant/vllm-glm53-flash-h100:v2-w4a16-dflash2 — vLLM nightly pin + baked drafter patches (EAGLE3 aux taps, drafter-aware KV partitioning); image of record for the 2026-09-28 8×H100 sweep
SM120 (RTX PRO 6000) cstechdev/vllm:glm53-flash-nope-sm120-cu130-20260826-r1
SM121 (DGX Spark) ghcr.io/canada-quant/vllm-glm53-flash-sm121:v2-w4a16-dflash2e, built from canada-quant/vllm-glm53-flash-sm121 with the two SM121 serving patches baked in

All configurations use expert parallelism. fp8 KV is not available on Hopper for this NoPE model.

H100 / H200, TP=4 (pinned image) — the Quick start command is the benchmarked recipe. num_speculative_tokens: 2 is the sweet spot on Hopper (52–55% acceptance; N=5 collapses acceptance to ~30%). Keep prefix caching off — it measured −2…−5% on H200. On 8× H200, two independent TP=4 replicas behind a load balancer beat one TP=8 endpoint by +28–38% aggregate at c128–c512; TP=8 wins single-stream and holds one 7.8M-token pool. On 2× H200 (89.5 GiB weights per GPU) add --tensor-parallel-size 2 --max-num-seqs 128 --max-cudagraph-capture-size 128 for 262K, or --max-model-len 1048576 --max-num-seqs 16 --gpu-memory-utilization 0.95 --max-cudagraph-capture-size 64 --max-num-batched-tokens 4096 for 1M.

8× H100 + DFlash2-G (v2 image, 2026-09-28 closeout) — latency recipe = TP8 + DFlash2-G K=7: c1 181.2 tok/s, +46% vs spec-off (124.1). Aggregate recipe = TP8 spec-off or MTP-N2: c256 2,418 / 2,330 tok/s; TP8 MTP-N2 is the balanced middle with the best grid c1 (215.0). Do not run DFlash2 at TP4 past c128 — the drafter's 8 standalone FullAttention KV tensors cut the KV pool ~6× and aggregate collapses (structural, confirmed uncontended: 694 tok/s @c128 vs TP8's 2,220); MTP-N2's shared-embedding draft softens the cliff (solo-TP4 c1 193.2, the best single-stream shape measured on this rig, c256 1,292). With any drafter on this image, cap --max-model-len ≤ 131071 — 128K-token prefills die on the speculative path (spec-off survives them). Full grids: BENCHMARKS.md. On the same rig and harness, NVIDIA's own NVFP4 checkpoint — an emulated FP4 path on Hopper (no native SM90 FP4 kernels; the engine warns at boot) — delivers 122.56 tok/s c1 / 2,299.84 @c256 at its best arm against our 181.2 / 2,418.1: on Hopper, use our quant + drafter + recipe (full head-to-head).

8× H100, TP=8, DFlash2-G (bench-validated 2026-09-28) — single node, 8192/1024 prompts: the latency recipe is TP=8 + DFlash2-G K=7: 181.2 tok/s c1, +46% vs spec-off 124.1; the aggregate recipe is TP=8 spec-off: 2,418 tok/s @c256 (MTP N=2 at 2,330 is the balanced middle and the best c1 at 215.0). K=4 beats K=7 on aggregate at every concurrency. Avoid TP=4 with the DFlash2 drafter past c32: its 8 standalone full-attention KV tensors cut the KV pool ~6× vs TP8 (311,999 vs 1,811,949 tokens) and force a preemption/recompute TTFT cliff that is structural, not contention — solo-TP4 measured 160.5 K=7 / 165.7 K=4 / 123.9 spec-off / 193.18 MTP-N2 tok/s c1, MTP-N2 being the best single-stream shape measured on this rig at any TP (1,292.4 @c256, the only speculative arm still rising there). With any drafter, cap --max-model-len ≤ 131071 — the speculative path faults on 131K-token prefills (spec-off survives); image ghcr.io/canada-quant/vllm-glm53-flash-h100:v2-w4a16-dflash2, cold boot ≈ 12 min.

RTX PRO 6000, TP=4 — same command with the SM120 image and --max-num-seqs 64 --max-num-batched-tokens 8192 --kv-cache-dtype fp8 --enable-prefix-caching. fp8 KV is required at 262K on 96 GB cards. Keep MTP on at every concurrency here: it adds +70% at c1 and +48% at c32.

2× DGX Spark, TP=2, 1M context — prebuilt image and one-command launcher in canada-quant/vllm-glm53-flash-sm121; the launcher also ships in the drafter repo. Drafter: canada-quant/GLM-5.3-Flash-DFlash2-G (the authors' self-trained DFlash2 drafter, Apache-2.0; 3.676 mean acceptance at K=7 on the 500-prompt holdout vs 3.632 for the incoai reference measured on the same hardware; drop-in successor of -F and -E, E having carried the 2026-09-22 banked Spark row and G being the Spark drafter of record since 2026-09-25 (the 2026-09-27 H2H champion leg and the live production serve both ran G, sha256-verified) — DRAFTER_HOST_PATH=/models/GLM-5.3-Flash-DFlash2-G). Start the worker rank first, wait 25 s, then the head rank.

# on both nodes, rank1 (worker) first, then rank0 (head) 25 s later
MAX_MODEL_LEN=1048576 KV_CACHE_MEM=9663676416 GMU=0.90 EAGER=0 GRAPHS=1 bash launch_dflash2_tp2.sh <rank>

Hard constraints (Spark): num_speculative_tokens must be 7 (any other count wedges boot); confirm the boot log shows the mask-embedding load (mask_token_id 154856); keep single prompts ≤ ~310K tokens; stop with docker stop -t 30, never rm -f. Cold boot is 6–10 minutes.

Quality

Benchmark Hardware W4A16 (this) Reference
AIME 2025 — n=120, max thinking, 131,072-token budget H100 0.8833 NVFP4 0.9000 — within noise (0.42σ)
AIME 2025 — same protocol RTX PRO 6000 0.8083 raw · 0.8833 with a budget-commit fix NVFP4 0.9000
AIME 2026 — n=120, max thinking, 131,072-token budget 2× DGX Spark 85.0% (102/120) EXL3 80.0% (96/120), matched protocol
GSM8K H100, RTX PRO 6000 0.970–0.975 parity across quants
GPQA-Diamond — n=198 @131K H100 · RTX PRO 6000 0.8586 · 0.8586 NVFP4 0.8687 · 0.8737 — within noise
MMMU (validation) — n=900, lmms-eval 0.7.3 task spec 8× H100 TP=8, spec-off, greedy 0.74667 (672/900) NVIDIA NVFP4 ckpt, same rig (2026-09-28): 0.68; first formal for this artifact; MC 0.766 · open 0.434
OCRBench — n=1000, lmms-eval 0.7.3 task spec 8× H100 TP=8, spec-off, greedy 888 /1000 NVIDIA NVFP4 ckpt, same rig (2026-09-28): 882; weakest: handwritten math 56/100; strongest: doc-VQA 192/200

On RTX PRO 6000, ≈63% of the raw AIME 2025 deficit is a budget wall (empty-answer rate 11.7–14.2% vs 3.3%) and ≈37% is SM120 kernel numerics; a zero-cost commit hook closes it but is not part of the published recipes. Vision (formal, 2026-09-28): MMMU validation 0.74667 (672/900) and OCRBench 888/1000 on the 8×H100 closeout stack (lmms-eval 0.7.3 task specs, greedy temp 0, seed 42). The vision tower is BF16 passthrough, not covered by the text-only calibration — these are the artifact's own vision baseline; the only same-rig vision baseline measured to date is the NVIDIA NVFP4 checkpoint (MMMU 0.68 / OCRBench 882, 2026-09-28). Details in BENCHMARKS.md.

Benchmarks summary

Output tok/s, thinking ON, same hardware, flags and prompts within each row.

Hardware W4A16 (this) Reference Read
8× H100 (SM90), p5 grid, 8192/1024 latency recipe TP8 + DFlash2-G K=7: c1 181.2 (+46% vs spec-off 124.1); aggregate recipe TP8 spec-off: c256 2,418; TP8 MTP-N2: c1 215.0 · c256 2,330 — (cross-rig; same-rig row below) solo-TP4: MTP-N2 c1 193.2 = best single-stream shape measured · c256 1,292; spec-off 1,549.9 @c256 (no cliff without the drafter); DFlash TP4 falls past c128
8× H100 vs NVIDIA NVFP4 ckpt (emulated FP4 path) — same rig, 2026-09-28 latency recipe c1 181.2 · aggregate c256 2,418 · best measured c1 193.18 (solo MTP-N2) NVFP4 N2 (closeout mirror, TP8): 122.56 · 2,299.84; N1b card recipe (TP4 × 2EP, KV bf16): 81.87 · 861.95, TTFT p50 263 s @c256 ours leads both headline cells (+47.8% c1, +5.1% c256); c128 dead heat 2,066.8 vs 2,081.13 (NVFP4 +0.7%); ckpt ships no MTP layer — full grid
8× H100 (same rig, identical harness) MTP-N2 TP4: c1 175.1 · c8 219.4 · c32 726.8; TPOT p50 2.99 ms NVFP4 MTP-N2 TP4: 166.5 · 239.2 · 864.8; TPOT 3.28 ms +5.2% c1 and TPOT −9% vs NVFP4; NVFP4 edges c8/c32 in this TP4 pool-limited cell — the TP8 rows above are the answer shape
4× RTX PRO 6000, TP=4, MTP N=2 109.9 · 318.5 · 534.4 NVFP4: 109.3 · 319.5 · 530.8 parity (±0.7%)
4× H200, TP=4, MTP N=2 c1 195 · c8 698 · c32 1,258 · c64 1,529 · c128 1,789 · c256 1,954 — 8× H200 TP=8: 217 → 2,681; two TP=4 replicas: 391 → 3,911 aggregate
8× H100, TP=8, DFlash2-G K=7 (latency recipe) c1 181.2 TP8 spec-off 124.1 +46% single-stream; aggregate recipe = TP8 spec-off c256 2,418; solo-TP4 MTP-N2 c1 193.18 is the best single-stream shape measured on this rig
2× DGX Spark, TP=2, DFlash2, 8K/256, 1M serve c1 33.0 · c2 35.6 · c4 59.6 · c6 67.1 EXL3 (matched protocol): 29.9 · 59.6 · 112.7; LibertAI NVFP4 ckpt (2026-08-31): did not boot (9/9 OOM) +10.4% c1 and +31–39% long-prefill vs EXL3; EXL3 leads mid-concurrency (c2 +67%, c4 +89%)
2× DGX Spark, TP=2, DFlash2-G @800k — H2H final 2026-09-27 tg32 d0 31.35±3.52 (30.88 on 09-25) · c1@65k 15.34 (10.72 on 09-25) · c1@100k 9.97 NVIDIA NVFP4 quant + serving stack w/ incoai DFlash2 drafter (legB): 34.82±1.92 · 29.19 · 17.95 NVFP4 stack ahead on all 28 cells (+16.9%…+254.5%), widest at depth; our one measured win = stability at depth (their run2 HTTP 500; ours clean ×2). Full grid: BENCHMARKS.md

vs NVIDIA's NVFP4 checkpoint — 8× H100, same rig (2026-09-28)

nvidia/GLM-5.3-Flash-NVFP4 is a Blackwell-only release and Hopper has no native FP4 math: on SM90 every weight dequantizes through FP4-emulation kernels (the engine warns at boot) — an emulated path, not a native comparison. Measured on the same p5 rig and closeout harness (8192/1024, thinking ON). The card-verbatim fp8-KV recipe cannot boot on Hopper (scale item not float32 in the KV-quant kernel); the card-recipe arm runs with KV bf16, the one documented deviation. Output tok/s:

runtime c1 c8 c32 c64 c128 c256
W4A16 TP8 + DFlash2-G K=7 (latency recipe) 181.2 633.0 1009.1 1471.4 1645.3 1653.5
W4A16 TP8 spec-off (aggregate recipe) 135.4 651.0 1375.9 1622.3 2066.8 2418.1
W4A16 TP4 MTP-N2 (solo, best single-stream shape) 193.18 559.2 1078.4 1181.9 1258.5 1292.4
NVFP4 N2 (closeout mirror, TP8 spec-off) 122.56 629.32 1052.71 1564.07 2081.13 2299.84
NVFP4 N1b (card recipe TP4 × 2EP, KV bf16) 81.87 185.01 577.28 823.02 845.24 861.95

Our recipes lead NVFP4's best arm (N2) by 47.8% at c1 (181.2 vs 122.56; solo MTP-N2 193.18 = +57.6%) and 5.1% at c256 (2,418.1 vs 2,299.84); c128 is a dead heat (2,066.8 vs 2,081.13, NVFP4 +0.7%). The card-recipe arm (N1b) collapses under load — TP4 KV exhaustion drives TTFT p50 to 263 s @c256 — and the checkpoint ships no MTP layer, so it cannot express GLM's native speculative decode at all. On Hopper, use our quant + drafter + recipe. Boot-failure autopsy and sha256 custody: BENCHMARKS.md.

Single-stream decode is insensitive to KV length up to ≥486K (RTX PRO 6000); at batch, long-KV decode plateaus at ~2–4 tok/s per stream and long prefills serialize at a ~6–8.5K tok/s aggregate ceiling. All grids, protocols and the H200 extended table: BENCHMARKS.md.

Hardware & context limits

Each row is the largest context serving-validated on that configuration.

Configuration GPUs Validated context Stack
DGX Spark GB10 (SM121), TP=2 2× 128 GB UMA 1M — KV pool 1,360,420 tokens (1.30× a full 1M request) DFlash2 drafter, fp8 KV
RTX PRO 6000 (SM120), TP=4 4× 96 GB 512K (486K prompts measured) MTP N=2, fp8 KV
H100 (SM90), TP=4 4× 80 GB 262K (256K prompts measured) MTP N=2, bf16 KV
H100 (SM90), TP=8 + DFlash2-G 8× 80 GB 131,071 with any drafter — 128K-token prefills die on the speculative path (spec-off survives 131,072); with MTP-N2 the 262K row above applies v2 image, K=7 latency / spec-off aggregate
H200 (SM90), TP=4 4× 141 GB 262K — KV pool 6.19M tokens (≈23 concurrent 262K requests) MTP N=2, bf16 KV
H200 (SM90), TP=8 8× 141 GB 262K — KV pool 7.79M tokens MTP N=2, bf16 KV
H200 (SM90), TP=2 2× 141 GB 1M — KV pool 2.84M tokens (929K-token prompt measured) MTP N=2, bf16 KV

Known issues

  • config.json (2026-09-08): vLLM matches quantization_config.ignore against its own fused module names, so the ignore list now carries both the HF and vLLM spellings plus re:.*\.layers\.45\..* for the MTP head. Older 765-entry copies fail at load. Weights unchanged.
  • DFlash2 admission wedge (SM90 research stack only, MTP recipes unaffected): with the DFlash2 drafter at block size 2304, prompts above ~15.5K tokens are never admitted. A fix was validated to 256K prompts; block size 1536 avoids it. Filed as vllm-project/vllm#55800.
  • Marlin no-split-K path on SM121: deterministic illegal memory access at M=256 when forcing split_k=1; the stock heuristic used in serving is clean. Filed as vllm-project/vllm#56064.

Quantization details

Field Value
Architecture Glm5NextForConditionalGeneration (glm5_next) — 45 decoder layers + MTP layer 45, 288 routed experts (top-8) + 1 shared, KDA + DSA attention, 24-block vision tower
Quantized 36,288 tensors = 42 MoE layers × 288 experts × 3 GEMMs — W4A16, INT4, symmetric, group 128, GPTQ, compressed-tensors pack-quantized
Kept in BF16 attention (incl. DSA indexer), dense prefix layers 0–2, shared experts, router, mHC tensors, embeddings, lm_head, norms, vision tower (348 keys), MTP layer 45 (889 keys)
Kept in FP32 A_log, dt_bias, e_score_correction_bias, hc_* — verbatim from source
Calibration 256 samples × 4096 tokens, in-distribution chat/code mix, sequential per-layer GPTQ
Built on 8× NVIDIA B300, 2026-08-27

Build gates, all passing: exactly 36,288 packed tensors and nothing quantized outside routed experts; vision key set 348/348 identical to source; MTP layer present; zero dtype drift vs source; no collapsed expert scales. Loads with transformers ≥ 5.16; text generation and image captioning smoke tests pass.

Citation & license

@misc{canada_quant_glm53_flash_w4a16_mtp,
  title  = {GLM-5.3-Flash W4A16 (INT4) + BF16 MTP},
  author = {canada-quant},
  year   = {2026},
  url    = {https://huggingface.co/canada-quant/GLM-5.3-Flash-W4A16-MTP}
}

MIT, inherited from the base model. Follow the base model's usage terms.


Built, benchmarked and documented with the Digby.ai coding harness, developed by CQL.ca.