GLM-5.3-Flash — W4A16 (INT4) + BF16 MTP
Model description
INT4 weight-only quantization of zai-org/GLM-5.3-Flash with the BF16 MTP draft head kept for speculative decoding. Only the 36,288 routed-expert GEMMs are INT4 (GPTQ, symmetric, group-size 128); attention, router, shared experts, embeddings, the vision tower and the MTP head stay in BF16. Not a new model — all capability comes from the base model.
| Size | 177.7 GiB (BF16 ≈ 599 GiB, −70%) |
| Runs on | NVIDIA H100, H200, RTX PRO 6000, and DGX Spark (GB10) — validated GPU counts and context per config in Hardware & context limits |
| Context | full 1,048,576 tokens on 2× DGX Spark and 2× H200; 262K–512K on the 4/8-GPU x86 configs (KV-memory-bound) |
| Quality | AIME 2025 0.8833 (n=120) vs 0.9000 for the NVFP4 reference on H100, within noise; GSM8K 0.97; GPQA-Diamond within noise; vision: MMMU 0.747 vs 0.68 NVFP4 (same rig) · OCRBench 888 vs 882 |
| Throughput | 8× H100: 181.2 tok/s single-stream latency recipe (TP8 + DFlash2-G K=7, +46%), 2,418 tok/s @c256 aggregate (TP8 spec-off); 4× H200 TP=4: 195 → 1,954 tok/s from c1 to c256; beats NVIDIA's NVFP4 checkpoint on H100 (+47.8% c1, +5.1% c256); matches the NVFP4 reference on RTX PRO 6000 |
Full benchmark grids, comparison protocols and research notes: BENCHMARKS.md.
Uses & recommended recipes
Quick start
# 1. Download (~178 GiB)
huggingface-cli download canada-quant/GLM-5.3-Flash-W4A16-MTP --local-dir /models/glm53-flash-w4a16-mtp
# 2. Serve on 4× H100 / H200 (other hardware: see Serving recipes)
docker run --gpus '"device=0,1,2,3"' --ipc=host --network=host --rm \
-v /models:/models vllm/vllm-openai:glm53-flash-x86_64-cu130 \
vllm serve /models/glm53-flash-w4a16-mtp --served-model-name glm53-w4 \
--tensor-parallel-size 4 --enable-expert-parallel \
--max-model-len 262144 --max-num-seqs 512 \
--gpu-memory-utilization 0.92 --no-enable-prefix-caching \
--speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
--reasoning-parser glm45 --tool-call-parser glm47 --enable-auto-tool-choice \
--trust-remote-code --port 8000
# 3. Call it (OpenAI-compatible)
curl http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "glm53-w4",
"messages": [{"role": "user", "content": "Prove there are infinitely many primes."}]
}'
Two things that bite. If your
config.jsonpredates 2026-09-08, re-download it — older copies fail in vLLM withKeyError: 'layers.0.mlp.gate_up_proj.weight'(weights are unchanged). And always pass--max-num-seqs ≤ 512— the vLLM default of 1024 exceeds this hybrid linear-attention model's 512 Mamba/KDA-state cache blocks.
Serving recipes
| Architecture | Image |
|---|---|
| SM90 (H100 / H200) | vllm/vllm-openai:glm53-flash-x86_64-cu130 (validated). Upstream vllm/vllm-openai:nightly-x86_64 ≥ 2026-09-08 also boots this checkpoint (vllm-project/vllm#53906); pass --attention-backend FLASH_ATTN_MLA_SPARSE there, its default backend faults on ≥131K prompts. |
| SM90, 8× H100 + DFlash2-G drafter | ghcr.io/canada-quant/vllm-glm53-flash-h100:v2-w4a16-dflash2 — vLLM nightly pin + baked drafter patches (EAGLE3 aux taps, drafter-aware KV partitioning); image of record for the 2026-09-28 8×H100 sweep |
| SM120 (RTX PRO 6000) | cstechdev/vllm:glm53-flash-nope-sm120-cu130-20260826-r1 |
| SM121 (DGX Spark) | ghcr.io/canada-quant/vllm-glm53-flash-sm121:v2-w4a16-dflash2e, built from canada-quant/vllm-glm53-flash-sm121 with the two SM121 serving patches baked in |
All configurations use expert parallelism. fp8 KV is not available on Hopper for this NoPE model.
H100 / H200, TP=4 (pinned image) — the Quick start command is the benchmarked recipe. num_speculative_tokens: 2 is the sweet spot on Hopper (52–55% acceptance; N=5 collapses acceptance to ~30%). Keep prefix caching off — it measured −2…−5% on H200. On 8× H200, two independent TP=4 replicas behind a load balancer beat one TP=8 endpoint by +28–38% aggregate at c128–c512; TP=8 wins single-stream and holds one 7.8M-token pool. On 2× H200 (89.5 GiB weights per GPU) add --tensor-parallel-size 2 --max-num-seqs 128 --max-cudagraph-capture-size 128 for 262K, or --max-model-len 1048576 --max-num-seqs 16 --gpu-memory-utilization 0.95 --max-cudagraph-capture-size 64 --max-num-batched-tokens 4096 for 1M.
8× H100 + DFlash2-G (v2 image, 2026-09-28 closeout) — latency recipe = TP8 + DFlash2-G K=7: c1 181.2 tok/s, +46% vs spec-off (124.1). Aggregate recipe = TP8 spec-off or MTP-N2: c256 2,418 / 2,330 tok/s; TP8 MTP-N2 is the balanced middle with the best grid c1 (215.0). Do not run DFlash2 at TP4 past c128 — the drafter's 8 standalone FullAttention KV tensors cut the KV pool ~6× and aggregate collapses (structural, confirmed uncontended: 694 tok/s @c128 vs TP8's 2,220); MTP-N2's shared-embedding draft softens the cliff (solo-TP4 c1 193.2, the best single-stream shape measured on this rig, c256 1,292). With any drafter on this image, cap --max-model-len ≤ 131071 — 128K-token prefills die on the speculative path (spec-off survives them). Full grids: BENCHMARKS.md. On the same rig and harness, NVIDIA's own NVFP4 checkpoint — an emulated FP4 path on Hopper (no native SM90 FP4 kernels; the engine warns at boot) — delivers 122.56 tok/s c1 / 2,299.84 @c256 at its best arm against our 181.2 / 2,418.1: on Hopper, use our quant + drafter + recipe (full head-to-head).
8× H100, TP=8, DFlash2-G (bench-validated 2026-09-28) — single node, 8192/1024 prompts: the latency recipe is TP=8 + DFlash2-G K=7: 181.2 tok/s c1, +46% vs spec-off 124.1; the aggregate recipe is TP=8 spec-off: 2,418 tok/s @c256 (MTP N=2 at 2,330 is the balanced middle and the best c1 at 215.0). K=4 beats K=7 on aggregate at every concurrency. Avoid TP=4 with the DFlash2 drafter past c32: its 8 standalone full-attention KV tensors cut the KV pool ~6× vs TP8 (311,999 vs 1,811,949 tokens) and force a preemption/recompute TTFT cliff that is structural, not contention — solo-TP4 measured 160.5 K=7 / 165.7 K=4 / 123.9 spec-off / 193.18 MTP-N2 tok/s c1, MTP-N2 being the best single-stream shape measured on this rig at any TP (1,292.4 @c256, the only speculative arm still rising there). With any drafter, cap --max-model-len ≤ 131071 — the speculative path faults on 131K-token prefills (spec-off survives); image ghcr.io/canada-quant/vllm-glm53-flash-h100:v2-w4a16-dflash2, cold boot ≈ 12 min.
RTX PRO 6000, TP=4 — same command with the SM120 image and --max-num-seqs 64 --max-num-batched-tokens 8192 --kv-cache-dtype fp8 --enable-prefix-caching. fp8 KV is required at 262K on 96 GB cards. Keep MTP on at every concurrency here: it adds +70% at c1 and +48% at c32.
2× DGX Spark, TP=2, 1M context — prebuilt image and one-command launcher in canada-quant/vllm-glm53-flash-sm121; the launcher also ships in the drafter repo. Drafter: canada-quant/GLM-5.3-Flash-DFlash2-G (the authors' self-trained DFlash2 drafter, Apache-2.0; 3.676 mean acceptance at K=7 on the 500-prompt holdout vs 3.632 for the incoai reference measured on the same hardware; drop-in successor of -F and -E, E having carried the 2026-09-22 banked Spark row and G being the Spark drafter of record since 2026-09-25 (the 2026-09-27 H2H champion leg and the live production serve both ran G, sha256-verified) — DRAFTER_HOST_PATH=/models/GLM-5.3-Flash-DFlash2-G). Start the worker rank first, wait 25 s, then the head rank.
# on both nodes, rank1 (worker) first, then rank0 (head) 25 s later
MAX_MODEL_LEN=1048576 KV_CACHE_MEM=9663676416 GMU=0.90 EAGER=0 GRAPHS=1 bash launch_dflash2_tp2.sh <rank>
Hard constraints (Spark): num_speculative_tokens must be 7 (any other count wedges boot); confirm the boot log shows the mask-embedding load (mask_token_id 154856); keep single prompts ≤ ~310K tokens; stop with docker stop -t 30, never rm -f. Cold boot is 6–10 minutes.
Quality
| Benchmark | Hardware | W4A16 (this) | Reference |
|---|---|---|---|
| AIME 2025 — n=120, max thinking, 131,072-token budget | H100 | 0.8833 | NVFP4 0.9000 — within noise (0.42σ) |
| AIME 2025 — same protocol | RTX PRO 6000 | 0.8083 raw · 0.8833 with a budget-commit fix | NVFP4 0.9000 |
| AIME 2026 — n=120, max thinking, 131,072-token budget | 2× DGX Spark | 85.0% (102/120) | EXL3 80.0% (96/120), matched protocol |
| GSM8K | H100, RTX PRO 6000 | 0.970–0.975 | parity across quants |
| GPQA-Diamond — n=198 @131K | H100 · RTX PRO 6000 | 0.8586 · 0.8586 | NVFP4 0.8687 · 0.8737 — within noise |
| MMMU (validation) — n=900, lmms-eval 0.7.3 task spec | 8× H100 TP=8, spec-off, greedy | 0.74667 (672/900) | NVIDIA NVFP4 ckpt, same rig (2026-09-28): 0.68; first formal for this artifact; MC 0.766 · open 0.434 |
| OCRBench — n=1000, lmms-eval 0.7.3 task spec | 8× H100 TP=8, spec-off, greedy | 888 /1000 | NVIDIA NVFP4 ckpt, same rig (2026-09-28): 882; weakest: handwritten math 56/100; strongest: doc-VQA 192/200 |
On RTX PRO 6000, ≈63% of the raw AIME 2025 deficit is a budget wall (empty-answer rate 11.7–14.2% vs 3.3%) and ≈37% is SM120 kernel numerics; a zero-cost commit hook closes it but is not part of the published recipes. Vision (formal, 2026-09-28): MMMU validation 0.74667 (672/900) and OCRBench 888/1000 on the 8×H100 closeout stack (lmms-eval 0.7.3 task specs, greedy temp 0, seed 42). The vision tower is BF16 passthrough, not covered by the text-only calibration — these are the artifact's own vision baseline; the only same-rig vision baseline measured to date is the NVIDIA NVFP4 checkpoint (MMMU 0.68 / OCRBench 882, 2026-09-28). Details in BENCHMARKS.md.
Benchmarks summary
Output tok/s, thinking ON, same hardware, flags and prompts within each row.
| Hardware | W4A16 (this) | Reference | Read |
|---|---|---|---|
| 8× H100 (SM90), p5 grid, 8192/1024 | latency recipe TP8 + DFlash2-G K=7: c1 181.2 (+46% vs spec-off 124.1); aggregate recipe TP8 spec-off: c256 2,418; TP8 MTP-N2: c1 215.0 · c256 2,330 | — (cross-rig; same-rig row below) | solo-TP4: MTP-N2 c1 193.2 = best single-stream shape measured · c256 1,292; spec-off 1,549.9 @c256 (no cliff without the drafter); DFlash TP4 falls past c128 |
| 8× H100 vs NVIDIA NVFP4 ckpt (emulated FP4 path) — same rig, 2026-09-28 | latency recipe c1 181.2 · aggregate c256 2,418 · best measured c1 193.18 (solo MTP-N2) | NVFP4 N2 (closeout mirror, TP8): 122.56 · 2,299.84; N1b card recipe (TP4 × 2EP, KV bf16): 81.87 · 861.95, TTFT p50 263 s @c256 | ours leads both headline cells (+47.8% c1, +5.1% c256); c128 dead heat 2,066.8 vs 2,081.13 (NVFP4 +0.7%); ckpt ships no MTP layer — full grid |
| 8× H100 (same rig, identical harness) | MTP-N2 TP4: c1 175.1 · c8 219.4 · c32 726.8; TPOT p50 2.99 ms | NVFP4 MTP-N2 TP4: 166.5 · 239.2 · 864.8; TPOT 3.28 ms | +5.2% c1 and TPOT −9% vs NVFP4; NVFP4 edges c8/c32 in this TP4 pool-limited cell — the TP8 rows above are the answer shape |
| 4× RTX PRO 6000, TP=4, MTP N=2 | 109.9 · 318.5 · 534.4 | NVFP4: 109.3 · 319.5 · 530.8 | parity (±0.7%) |
| 4× H200, TP=4, MTP N=2 | c1 195 · c8 698 · c32 1,258 · c64 1,529 · c128 1,789 · c256 1,954 | — | 8× H200 TP=8: 217 → 2,681; two TP=4 replicas: 391 → 3,911 aggregate |
| 8× H100, TP=8, DFlash2-G K=7 (latency recipe) | c1 181.2 | TP8 spec-off 124.1 | +46% single-stream; aggregate recipe = TP8 spec-off c256 2,418; solo-TP4 MTP-N2 c1 193.18 is the best single-stream shape measured on this rig |
| 2× DGX Spark, TP=2, DFlash2, 8K/256, 1M serve | c1 33.0 · c2 35.6 · c4 59.6 · c6 67.1 | EXL3 (matched protocol): 29.9 · 59.6 · 112.7; LibertAI NVFP4 ckpt (2026-08-31): did not boot (9/9 OOM) | +10.4% c1 and +31–39% long-prefill vs EXL3; EXL3 leads mid-concurrency (c2 +67%, c4 +89%) |
| 2× DGX Spark, TP=2, DFlash2-G @800k — H2H final 2026-09-27 | tg32 d0 31.35±3.52 (30.88 on 09-25) · c1@65k 15.34 (10.72 on 09-25) · c1@100k 9.97 | NVIDIA NVFP4 quant + serving stack w/ incoai DFlash2 drafter (legB): 34.82±1.92 · 29.19 · 17.95 | NVFP4 stack ahead on all 28 cells (+16.9%…+254.5%), widest at depth; our one measured win = stability at depth (their run2 HTTP 500; ours clean ×2). Full grid: BENCHMARKS.md |
vs NVIDIA's NVFP4 checkpoint — 8× H100, same rig (2026-09-28)
nvidia/GLM-5.3-Flash-NVFP4 is a Blackwell-only release and Hopper has no native FP4 math: on SM90 every weight dequantizes through FP4-emulation kernels (the engine warns at boot) — an emulated path, not a native comparison. Measured on the same p5 rig and closeout harness (8192/1024, thinking ON). The card-verbatim fp8-KV recipe cannot boot on Hopper (scale item not float32 in the KV-quant kernel); the card-recipe arm runs with KV bf16, the one documented deviation. Output tok/s:
| runtime | c1 | c8 | c32 | c64 | c128 | c256 |
|---|---|---|---|---|---|---|
| W4A16 TP8 + DFlash2-G K=7 (latency recipe) | 181.2 | 633.0 | 1009.1 | 1471.4 | 1645.3 | 1653.5 |
| W4A16 TP8 spec-off (aggregate recipe) | 135.4 | 651.0 | 1375.9 | 1622.3 | 2066.8 | 2418.1 |
| W4A16 TP4 MTP-N2 (solo, best single-stream shape) | 193.18 | 559.2 | 1078.4 | 1181.9 | 1258.5 | 1292.4 |
| NVFP4 N2 (closeout mirror, TP8 spec-off) | 122.56 | 629.32 | 1052.71 | 1564.07 | 2081.13 | 2299.84 |
| NVFP4 N1b (card recipe TP4 × 2EP, KV bf16) | 81.87 | 185.01 | 577.28 | 823.02 | 845.24 | 861.95 |
Our recipes lead NVFP4's best arm (N2) by 47.8% at c1 (181.2 vs 122.56; solo MTP-N2 193.18 = +57.6%) and 5.1% at c256 (2,418.1 vs 2,299.84); c128 is a dead heat (2,066.8 vs 2,081.13, NVFP4 +0.7%). The card-recipe arm (N1b) collapses under load — TP4 KV exhaustion drives TTFT p50 to 263 s @c256 — and the checkpoint ships no MTP layer, so it cannot express GLM's native speculative decode at all. On Hopper, use our quant + drafter + recipe. Boot-failure autopsy and sha256 custody: BENCHMARKS.md.
Single-stream decode is insensitive to KV length up to ≥486K (RTX PRO 6000); at batch, long-KV decode plateaus at ~2–4 tok/s per stream and long prefills serialize at a ~6–8.5K tok/s aggregate ceiling. All grids, protocols and the H200 extended table: BENCHMARKS.md.
Hardware & context limits
Each row is the largest context serving-validated on that configuration.
| Configuration | GPUs | Validated context | Stack |
|---|---|---|---|
| DGX Spark GB10 (SM121), TP=2 | 2× 128 GB UMA | 1M — KV pool 1,360,420 tokens (1.30× a full 1M request) | DFlash2 drafter, fp8 KV |
| RTX PRO 6000 (SM120), TP=4 | 4× 96 GB | 512K (486K prompts measured) | MTP N=2, fp8 KV |
| H100 (SM90), TP=4 | 4× 80 GB | 262K (256K prompts measured) | MTP N=2, bf16 KV |
| H100 (SM90), TP=8 + DFlash2-G | 8× 80 GB | 131,071 with any drafter — 128K-token prefills die on the speculative path (spec-off survives 131,072); with MTP-N2 the 262K row above applies | v2 image, K=7 latency / spec-off aggregate |
| H200 (SM90), TP=4 | 4× 141 GB | 262K — KV pool 6.19M tokens (≈23 concurrent 262K requests) | MTP N=2, bf16 KV |
| H200 (SM90), TP=8 | 8× 141 GB | 262K — KV pool 7.79M tokens | MTP N=2, bf16 KV |
| H200 (SM90), TP=2 | 2× 141 GB | 1M — KV pool 2.84M tokens (929K-token prompt measured) | MTP N=2, bf16 KV |
Known issues
config.json(2026-09-08): vLLM matchesquantization_config.ignoreagainst its own fused module names, so the ignore list now carries both the HF and vLLM spellings plusre:.*\.layers\.45\..*for the MTP head. Older 765-entry copies fail at load. Weights unchanged.- DFlash2 admission wedge (SM90 research stack only, MTP recipes unaffected): with the DFlash2 drafter at block size 2304, prompts above ~15.5K tokens are never admitted. A fix was validated to 256K prompts; block size 1536 avoids it. Filed as vllm-project/vllm#55800.
- Marlin no-split-K path on SM121: deterministic illegal memory access at M=256 when forcing
split_k=1; the stock heuristic used in serving is clean. Filed as vllm-project/vllm#56064.
Quantization details
| Field | Value |
|---|---|
| Architecture | Glm5NextForConditionalGeneration (glm5_next) — 45 decoder layers + MTP layer 45, 288 routed experts (top-8) + 1 shared, KDA + DSA attention, 24-block vision tower |
| Quantized | 36,288 tensors = 42 MoE layers × 288 experts × 3 GEMMs — W4A16, INT4, symmetric, group 128, GPTQ, compressed-tensors pack-quantized |
| Kept in BF16 | attention (incl. DSA indexer), dense prefix layers 0–2, shared experts, router, mHC tensors, embeddings, lm_head, norms, vision tower (348 keys), MTP layer 45 (889 keys) |
| Kept in FP32 | A_log, dt_bias, e_score_correction_bias, hc_* — verbatim from source |
| Calibration | 256 samples × 4096 tokens, in-distribution chat/code mix, sequential per-layer GPTQ |
| Built on | 8× NVIDIA B300, 2026-08-27 |
Build gates, all passing: exactly 36,288 packed tensors and nothing quantized outside routed experts; vision key set 348/348 identical to source; MTP layer present; zero dtype drift vs source; no collapsed expert scales. Loads with transformers ≥ 5.16; text generation and image captioning smoke tests pass.
Citation & license
@misc{canada_quant_glm53_flash_w4a16_mtp,
title = {GLM-5.3-Flash W4A16 (INT4) + BF16 MTP},
author = {canada-quant},
year = {2026},
url = {https://huggingface.co/canada-quant/GLM-5.3-Flash-W4A16-MTP}
}
MIT, inherited from the base model. Follow the base model's usage terms.
Built, benchmarked and documented with the Digby.ai coding harness, developed by CQL.ca.