keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4
Abliterated Vontra GLM-5.3-Flash MLX oQ4 for Mac Studio oMLX 0.6.3rc2 Dual-ANE. Same Dealign o_proj L15–45 transplant as the Spark NVFP4 pack — L0–14 stay stock, MTP included.
See RESPONSIBLE_USE.md and the gate form above. Access is gated with automatic approval after you agree.
| HF (these weights) | https://huggingface.co/drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4 |
| GitHub (oMLX Dual-ANE recipe) | https://github.com/drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4 |
| Spark NVFP4 cousin (same ablit) | HF · GitHub |
| Stock oQ4 parent | Vontra/GLM-5.3-Flash-MLX-oQ4-MTP |
Ablit source (o_proj L15–45) |
dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4 |
| Upstream | zai-org/GLM-5.3-Flash |
| Ablit | L15–45 self_attn.o_proj as BF16 (31 tensors, includes MTP) · L0–14 affine oQ4 stock |
| Gate | 32/32 bypass, 0 refuse, 0 garble |
| Decode (M3 Ultra 256 GB, oMLX 0.6.3rc2) | ~24 tok/s (192-token gens 22.7–25.0; 128-token 23.4 / 23.9). Stock oQ4 on the same box: 26.4 tok/s. |
These weights have safety refusals removed. Research / red-team only — you supply the guardrails.
What changed vs Vontra oQ4
Blackfrost-style rank-1 projection + affine requant stayed at 19/32 on this quant. The Spark 32/32 dest is a byte-copy of Dealign o_proj. On MLX we do not requantize those tensors: they land as BF16 nn.Linear weights (model-ablit-oproj-l15-45.safetensors). Experts, vision, QKV, and L0–14 o_proj stay Vontra affine oQ4.
One-shot (Mac Studio)
git clone https://github.com/drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4.git
cd keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4
# Hugging Face gate: agree on the model card, then `hf auth login`
bash oneshot-setup.sh
omlx serve --model-dir ~/.omlx/models --host 0.0.0.0 --port 11500
Serve id: glm53-flash-oq4-mtp-ablit-l15-45. Dual-ANE tile 4096, Lightning MTP on, thinking off.
Credits
Cite the original authors first.
| Who | What we used |
|---|---|
| Z.ai | Upstream GLM-5.3-Flash |
| Vontra | Mac oQ4 body; L0–14 o_proj, experts, vision stay theirs |
| dealignai (compute: Jordan Schenck) | BF16 o_proj L15–45 + MTP that reaches 32/32 |
| jundot | oMLX, oQ, Lightning MTP, Dual-ANE |
| onthehub97 | Dual-ANE/GPU prompt processing |
| Blaizzy | day-0 glm5_next in mlx-vlm |
| PipeNetwork | ClampedSwiGLU / Mac glm5_next runtime |
| ml-explore | MLX |
| Blackfrost | Direction reference; not this checkpoint |
| LibertAI | Spark NVFP4 layout Dealign o_proj was extracted from |
Pack: drowzeys / keys. Donate: GoFundMe.