drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4

🤗 Hugging Face 来源image-text-to-textmit321B 参数643 GBsafetensors✓ 37 个校验和今天更新
帮助这个模型通过 Pirate Face 分发

模型卡、文件列表和校验和已在此收录。如果你持有文件并有权分享,可以提交种子,让其他人从节点下载。

获取来源文件需要 Hugging Face 批准。
为此模型做种

keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4

Abliterated Vontra GLM-5.3-Flash MLX oQ4 for Mac Studio oMLX 0.6.3rc2 Dual-ANE. Same Dealign o_proj L15–45 transplant as the Spark NVFP4 pack — L0–14 stay stock, MTP included.

See RESPONSIBLE_USE.md and the gate form above. Access is gated with automatic approval after you agree.

HF (these weights) https://huggingface.co/drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4
GitHub (oMLX Dual-ANE recipe) https://github.com/drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4
Spark NVFP4 cousin (same ablit) HF · GitHub
Stock oQ4 parent Vontra/GLM-5.3-Flash-MLX-oQ4-MTP
Ablit source (o_proj L15–45) dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4
Upstream zai-org/GLM-5.3-Flash
Ablit L15–45 self_attn.o_proj as BF16 (31 tensors, includes MTP) · L0–14 affine oQ4 stock
Gate 32/32 bypass, 0 refuse, 0 garble
Decode (M3 Ultra 256 GB, oMLX 0.6.3rc2) ~24 tok/s (192-token gens 22.7–25.0; 128-token 23.4 / 23.9). Stock oQ4 on the same box: 26.4 tok/s.

These weights have safety refusals removed. Research / red-team only — you supply the guardrails.

What changed vs Vontra oQ4

Blackfrost-style rank-1 projection + affine requant stayed at 19/32 on this quant. The Spark 32/32 dest is a byte-copy of Dealign o_proj. On MLX we do not requantize those tensors: they land as BF16 nn.Linear weights (model-ablit-oproj-l15-45.safetensors). Experts, vision, QKV, and L0–14 o_proj stay Vontra affine oQ4.

One-shot (Mac Studio)

git clone https://github.com/drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4.git
cd keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4
# Hugging Face gate: agree on the model card, then `hf auth login`
bash oneshot-setup.sh
omlx serve --model-dir ~/.omlx/models --host 0.0.0.0 --port 11500

Serve id: glm53-flash-oq4-mtp-ablit-l15-45. Dual-ANE tile 4096, Lightning MTP on, thinking off.

Credits

Cite the original authors first.

Who What we used
Z.ai Upstream GLM-5.3-Flash
Vontra Mac oQ4 body; L0–14 o_proj, experts, vision stay theirs
dealignai (compute: Jordan Schenck) BF16 o_proj L15–45 + MTP that reaches 32/32
jundot oMLX, oQ, Lightning MTP, Dual-ANE
onthehub97 Dual-ANE/GPU prompt processing
Blaizzy day-0 glm5_next in mlx-vlm
PipeNetwork ClampedSwiGLU / Mac glm5_next runtime
ml-explore MLX
Blackfrost Direction reference; not this checkpoint
LibertAI Spark NVFP4 layout Dealign o_proj was extracted from

Pack: drowzeys / keys. Donate: GoFundMe.