pipenetwork/GLM-5.2-MLX-nvfp4

🤗 Hugging Face 来源text-generationmit743B 参数1.5 TBsafetensors✓ 92 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo pipenetwork/GLM-5.2-MLX-nvfp4 ./model-folder
需要做种者 →

GLM-5.2-MLX-nvfp4

Runtime — updated 2026-08-28: load with --trust-remote-code

This repository now bundles glm_moe_dsa.py (declared via model_file in config.json), a fixed runtime for this architecture, and needs it:

mlx_lm.generate --model pipenetwork/GLM-5.2-MLX-nvfp4 --trust-remote-code --prompt "..." --max-tokens 300

mlx-lm's own glm_moe_dsa builds a lightning indexer on all 78 layers, but GLM-5.2 ships indexer weights on 21 (indexer_types: the other 57 "shared" layers reuse the previous full layer's top-k selection). mlx_lm.load loads leniently and left those 57 indexers at random initialisation. Prompts up to 2048 tokens were unaffected (the indexer is bypassed below index_topk); beyond that, 57 of 78 layers attended to keys chosen by random projections. The bundled runtime implements the schedule as the reference does (plus fp32 indexer scores and router logits and the indexer LayerNorm epsilon); tiny-config parity against transformers 5.16 is 4e-7 with the sparse path live, and a strict load of this checkpoint reports zero missing and zero unexpected tensors. Details, tests and the GLM-5.3 builds made with it: github.com/PipeNetwork/glm53-mlx. The weights are unchanged.

An MLX conversion of zai-org/GLM-5.2 quantized to NVFP4 (4-bit FP4, group size 16) for Apple Silicon with mlx-lm.

This is the MLX analog of NVIDIA's nvidia/GLM-5.2-NVFP4. NVIDIA's checkpoint stores weights in ModelOpt-packed NVFP4 that mlx-lm cannot read directly, so this build was produced by quantizing the bf16 base with MLX's own NVFP4 mode (--q-mode nvfp4 --q-group-size 16).

  • Base model: zai-org/GLM-5.2 (GlmMoeDsaForCausalLM, 753B total / ~40B active MoE, text-only)
  • Format: MLX, NVFP4 (4-bit FP4, group size 16)
  • Approx. size on disk: 390G
  • Converted with: mlx-lm 0.31.2

Usage

pip install -U mlx-lm
mlx_lm.generate --model pipenetwork/GLM-5.2-MLX-nvfp4 --prompt "Explain mixture-of-experts in one sentence." --max-tokens 128

License

MIT, inherited from the base model.