pipenetwork/GLM-5.2-REAP50-MLX-4bit

🤗 Hugging Face 来源text-generationmit381B 参数762 GBsafetensors✓ 47 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo pipenetwork/GLM-5.2-REAP50-MLX-4bit ./model-folder
需要做种者 →

GLM-5.2-REAP50-MLX-4bit

Runtime — updated 2026-08-28: load with --trust-remote-code

This repository now bundles glm_moe_dsa.py (declared via model_file in config.json), a fixed runtime for this architecture, and needs it:

mlx_lm.generate --model pipenetwork/GLM-5.2-REAP50-MLX-4bit --trust-remote-code --prompt "..." --max-tokens 300

mlx-lm's own glm_moe_dsa builds a lightning indexer on all 78 layers, but GLM-5.2 ships indexer weights on 21 (indexer_types: the other 57 "shared" layers reuse the previous full layer's top-k selection). mlx_lm.load loads leniently and left those 57 indexers at random initialisation. Prompts up to 2048 tokens were unaffected (the indexer is bypassed below index_topk); beyond that, 57 of 78 layers attended to keys chosen by random projections. The bundled runtime implements the schedule as the reference does (plus fp32 indexer scores and router logits and the indexer LayerNorm epsilon); tiny-config parity against transformers 5.16 is 4e-7 with the sparse path live, and a strict load of this checkpoint reports zero missing and zero unexpected tensors. Details, tests and the GLM-5.3 builds made with it: github.com/PipeNetwork/glm53-mlx. The weights are unchanged.

REAP expert-pruned + 4-bit MLX conversion of zai-org/GLM-5.2. Keeps the 128 most-salient experts per layer (of 256) → ~394B params, smaller/faster than the full model.

⚠️ Quality warning: this is the most aggressive prune (50% of experts). Held-out perplexity is +37.5% vs the full model — visibly degraded. Use only if you need the smallest/fastest footprint and have validated it for your task.

What is this

Pruned with REAP (Router-weighted Expert Activation Pruning, Cerebras / ICLR 2026): per MoE layer, experts are scored by mean(router_gate_weight × ‖expert_output‖) over a calibration set; the lowest-saliency experts are dropped and the router is sliced to the survivors. No retraining. n_routed_experts reduced 256→128.

Quality (held-out perplexity, Frankenstein — not in calibration)

Variant Experts ~Params Held-out PPL vs full
full GLM-5.2 (4-bit) 256 ~750B 1.447 —
REAP25 192 ~572B 1.481 +2.3%
REAP37 160 ~480B 1.553 +7.3%
REAP50 (this repo) 128 ~394B 1.990 +37.5%

This variant: PPL 1.990 (+37.5% vs full) — SMALLEST but notably degraded. (Absolute PPL is low because the eval text is highly predictable; treat the numbers as relative degradation.)

Methodology

Calibrated on the 4-bit GLM-5.2 (192 seqs × 1024 tok, prose + code); pruned during MLX conversion (no intermediate bf16). Requires the glm_moe_dsa / deepseek_v32 MLX path with per-layer indexer handling.

Use with mlx-lm

pip install mlx-lm
python -m mlx_lm generate --model pipenetwork/GLM-5.2-REAP50-MLX-4bit --prompt "Hello" -m 256

License

MIT (inherited from GLM-5.2). Quantization: {"group_size": 64, "bits": 4, "mode": "affine"}.