RadixArk/glm47-flash-blockwise-fp8

🤗 Hugging Face sourcetext-generationmit31.2B params33 GBsafetensors✓ 49 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo RadixArk/glm47-flash-blockwise-fp8 ./model-folder
Needs a seeder →

GLM-4.7-Flash blockwise FP8

Blockwise FP8 quantization of zai-org/GLM-4.7-Flash: e4m3 weights with one scale per 128×128 block and dynamic activation scaling, about half the size of the BF16 original (30 GiB vs 58 GiB). See the original model card for model details.

Quantization

  • FP8: the linear layers of attention, the dense MLP, and all routed and shared experts, including the MTP layer.
  • BF16: embeddings, lm_head, norms, and the MoE router.

Evaluation

gsm8k (5-shot, full test set) on SGLang with MTP speculative decoding: 0.809 for FP8 vs 0.819 for BF16; average accept length 2.41 vs 2.43.

Usage (SGLang)

python -m sglang.launch_server --model-path RadixArk/glm47-flash-blockwise-fp8 --tp-size 2 \
  --speculative-algorithm EAGLE --speculative-num-steps 2 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 3

Attention tensor parallelism must be 1 or 2: at 4, each rank's kv_b_proj shard (2240 rows) is not a multiple of the 128-row block. For more GPUs, use DP attention with expert parallelism, for example --tp-size 4 --dp-size 4 --enable-dp-attention --ep-size 4 --moe-a2a-backend deepep --cuda-graph-max-bs-decode 128.

License

MIT, same as the original model by Z.ai.