GLM-4.7-Flash blockwise FP8
Blockwise FP8 quantization of zai-org/GLM-4.7-Flash: e4m3 weights with one scale per 128×128 block and dynamic activation scaling, about half the size of the BF16 original (30 GiB vs 58 GiB). See the original model card for model details.
Quantization
- FP8: the linear layers of attention, the dense MLP, and all routed and shared experts, including the MTP layer.
- BF16: embeddings,
lm_head, norms, and the MoE router.
Evaluation
gsm8k (5-shot, full test set) on SGLang with MTP speculative decoding: 0.809 for FP8 vs 0.819 for BF16; average accept length 2.41 vs 2.43.
Usage (SGLang)
python -m sglang.launch_server --model-path RadixArk/glm47-flash-blockwise-fp8 --tp-size 2 \
--speculative-algorithm EAGLE --speculative-num-steps 2 \
--speculative-eagle-topk 1 --speculative-num-draft-tokens 3
Attention tensor parallelism must be 1 or 2: at 4, each rank's kv_b_proj shard (2240 rows) is not a multiple of the 128-row block. For more GPUs, use DP attention with expert parallelism, for example --tp-size 4 --dp-size 4 --enable-dp-attention --ep-size 4 --moe-a2a-backend deepep --cuda-graph-max-bs-decode 128.
License
MIT, same as the original model by Z.ai.