keys-GLM-5.3-Flash-NVFP4-ablit-l15-43-mtp-l45
Update 9/27/26 — compatible with NVIDIA GLM-5.3-Flash NVFP4. Same Keys transplant applies to both
nvidia/GLM-5.3-Flash-NVFP4andRedHatAI/GLM-5.3-Flash-NVFP4. NVIDIA apply: scripts/APPLY-NVIDIA.md.
Altered Dealign abliteration, anchor-safe, adapted for DGX Spark. Ablit source is Dealign o_proj with reference layers left stock for anchor protection (L0–14 and L44). Edited: L15–43 + MTP L45 (30 tensors). Published HF files are the RedHat body; NVIDIA official NVFP4 is an eligible parent for the same poke.
GitHub (method + scripts): drowzeys/keys-GLM-5.3-Flash-NVFP4-ablit-l15-43-mtp-l45 · METHOD.md.
See RESPONSIBLE_USE.md and the gate form above. Access is gated with automatic approval after you agree.
| HF | https://huggingface.co/drowzeys/keys-GLM-5.3-Flash-NVFP4-ablit-l15-43-mtp-l45 |
| GitHub (method + scripts) | https://github.com/drowzeys/keys-GLM-5.3-Flash-NVFP4-ablit-l15-43-mtp-l45 |
| Keys 0731 ancestor | HF anchored-tensors · GitHub 1M recipe |
| Parents | RedHatAI/GLM-5.3-Flash-NVFP4 (this dest) · nvidia/GLM-5.3-Flash-NVFP4 (eligible 9/27/26) |
| Ablit source | Altered Dealign from dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4 |
| Reference / anchor layers | L0–14 stock · L44 stock |
| Edited | L15–43 + MTP L45 self_attn.o_proj · 30 tensors |
| Upstream | zai-org/GLM-5.3-Flash |
| Gate | 32/32 bypass, 0 refuse, 0 garble (NVIDIA parent also 32/32 + 22/22 cyber, 9/27/26) |
| Layout (this repo) | 10 shards + model_mtp.safetensors, quant_method=compressed-tensors |
Preferred method: 0731 safety-anchors (enhance, don't hinder)
Keys learned this on DeepSeek-V4-Flash 0731. Projecting residual writes through early layers made the target stop refusing and made the stock drafter keep proposing refusal-shaped tokens.
Spare early layers. That is still our default. On GLM-5.3-Flash the in-checkpoint drafter is MTP layers.45, and residual refusal lived late + MTP.
This dest is an altered Dealign: we keep L0–14 and L44 as reference layers for anchor protection, and we transplant L15–43 + MTP L45. Dealign L44 is Δrel 0.74 (garble risk). The previous LibertAI ModelOpt publish included L44; this one does not.
As of 9/27/26 the same poke is compatible with NVIDIA GLM-5.3-Flash NVFP4 (attention stays BF16). Abliteration does not change FLOPs. The intended win is direct completions instead of refuse/hedge loops.
Credit: RedHat and NVIDIA (eligible NVFP4 parents)
RedHatAI/GLM-5.3-Flash-NVFP4 is the published parent in this repo: compressed-tensors NVFP4, ~193 GiB, 10 shards + MTP. Experts, vision, QKV, embeddings, L0–14 o_proj, and L44 o_proj remain theirs.
nvidia/GLM-5.3-Flash-NVFP4 is an eligible parent as of 9/27/26. 33-shard ModelOpt NVFP4 with every self_attn block in BF16. Same 30-tensor poke; dest stays NVIDIA-shape. Apply: APPLY-NVIDIA.md.
Credit: Dealign (altered; reference layers for anchor protection)
Full credit to dealignai / @dealignai (compute @jordanschenck) for GLM-5.3-Flash-UNCENSORED-NVFP4.
This dest is an altered Dealign: we byte-copy BF16 self_attn.o_proj for L15–43 and MTP L45. L0–14 and L44 stay parent stock as reference layers for anchor protection. We did not ship their full checkpoint as a swap.
Credit: Z.ai and the Spark vLLM recipe
- Z.ai / zai-org — GLM-5.3-Flash (
Glm5NextForConditionalGeneration, hybrid KDA+DSA, mHC, native MTP). - vLLM PR #53906.
- tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark — SM121 DFlash2 image,
--block-size 2304; barrydeen GMU 0.85 floor.
Abliteration recipe (published)
Altered Dealign: byte-copy o_proj into a RedHat or NVIDIA NVFP4 tree. Offsets from that tree's headers — never the LibertAI 120-shard map.
| Tensor | model.language_model.layers.{L}.self_attn.o_proj.weight (BF16) |
| Layers | 15–43 and 45 (30 tensors, includes MTP layers.45) |
| Skip | L44 RedHat stock (Dealign L44 Δrel 0.74) |
| Safety | L0–14 byte-identical to RedHat stock |
| Experts | NVFP4 passthrough (even in rewritten shards) |
| Gate | 32/32 bypass, 0 refuse, 0 garble, raw vLLM + thinking-off template |
Variation table: GitHub METHOD.md. Artifacts: ABLIT_META.json, VARIATIONS.json.
Reproduce:
# RedHat (this repo's published layout)
python3 scripts/apply_oproj_l15_45.py \
--src /path/to/GLM-5.3-Flash-NVFP4-RedHat \
--dst /path/to/dest \
--bins ./oproj_bins \
--skip-layers 44 \
--fresh
# NVIDIA official NVFP4 — flatten HF-cache first; see APPLY-NVIDIA.md
python3 scripts/apply_oproj_l15_45.py \
--src /path/to/nvidia-GLM-5.3-Flash-NVFP4-serve \
--dst /path/to/nvidia-GLM-5.3-Flash-NVFP4-serve \
--bins ./oproj_bins \
--skip-layers 44 \
--in-place
Thinking leak (template, not ablit)
Stock GLM-5.3-Flash always opens <think> and injects Reasoning Effort: Max. enable_thinking=false is a silent no-op — CoT lands in content. Mount chat_template.thinking-off.jinja over chat_template.jinja at serve time. Stock chat_template.jinja in this repo is unchanged from RedHat.
Files
| Path | Purpose |
|---|---|
model-00001-of-00010.safetensors … 00010 + model_mtp.safetensors + index |
Full NVFP4 checkpoint (RedHat layout, L15–43+L45 o_proj from Dealign) |
ABLIT_META.json |
Edit stats / recipe fingerprint |
VARIATIONS.json |
Refusal32 table |
chat_template.thinking-off.jinja |
Recommended serve overlay (closed <think></think>) |
tokenizer.json / chat_template.jinja / processor |
Unchanged from RedHat / Z.ai |
config.json is RedHat stock (index_topk 2048, quant_method=compressed-tensors). Honor it — do not pass --quantization modelopt_fp4.
The previous 120-shard LibertAI ModelOpt files (model-*-of-00120.safetensors) are removed from this repo.
Download
# after you agree to the gate (automatic approval)
hf download drowzeys/keys-GLM-5.3-Flash-NVFP4-ablit-l15-43-mtp-l45 \
--local-dir ~/models/GLM-5.3-Flash-NVFP4-ablit-l15-43-mtp-l45
Serve with Tony’s 2× DGX Spark DFlash2 recipe (marlin MoE, DFlash2 k=7, fp8 KV, --block-size 2304). GPU memory utilization ≤ 0.85. Mount the thinking-off template.
License
MIT, inherited from Z.AI GLM-5.3-Flash (also the RedHat NVFP4 card). You must still comply with the Responsible Use gate above.