RadixArk/glm47-flash-blockwise-fp8

🤗 Hugging Face 来源text-generationmit31.2B 参数33 GBsafetensors✓ 49 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo RadixArk/glm47-flash-blockwise-fp8 ./model-folder
需要做种者 →

GLM-4.7-Flash blockwise FP8

Blockwise FP8 quantization of zai-org/GLM-4.7-Flash: e4m3 weights with one scale per 128×128 block and dynamic activation scaling, about half the size of the BF16 original (30 GiB vs 58 GiB). See the original model card for model details.

Quantization

  • FP8: the linear layers of attention, the dense MLP, and all routed and shared experts, including the MTP layer.
  • BF16: embeddings, lm_head, norms, and the MoE router.

Evaluation

gsm8k (5-shot, full test set) on SGLang with MTP speculative decoding: 0.809 for FP8 vs 0.819 for BF16; average accept length 2.41 vs 2.43.

Usage (SGLang)

python -m sglang.launch_server --model-path RadixArk/glm47-flash-blockwise-fp8 --tp-size 2 \
  --speculative-algorithm EAGLE --speculative-num-steps 2 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 3

Attention tensor parallelism must be 1 or 2: at 4, each rank's kv_b_proj shard (2240 rows) is not a multiple of the 128-row block. For more GPUs, use DP attention with expert parallelism, for example --tp-size 4 --dp-size 4 --enable-dp-attention --ep-size 4 --moe-a2a-backend deepep --cuda-graph-max-bs-decode 128.

License

MIT, same as the original model by Z.ai.