GLM-5.3-Flash-Uncensored — EXL3 K2 (~2.0 bpw experts)
Sub-4-bit EXL3 (QTIP trellis) quantization of
orcarouter/GLM-5.3-Flash-Uncensored-FP8,
an abliterated finetune of zai-org/GLM-5.3-Flash
(glm5_next / Glm5NextForConditionalGeneration). To our knowledge the first
sub-4-bit EXL3 release of any uncensored GLM-5.3-Flash variant.
Sized for single-device serving: all 288 routed experts per MoE layer kept (no pruning), vision tower and NextN/MTP head kept, everything non-expert in native BF16.
| Slab | Format | Approx. size |
|---|---|---|
| Routed experts (43 MoE layers × 288 experts × 3 projections = 37,152 matrices) | EXL3 trellis, K=2 (~2.01 bpw payload) | ~73 GiB |
Attention / KDA / shared experts / routers / embeddings / lm_head / vision / MTP |
BF16 (native) | ~18 GiB |
| Total weights | ~91 GiB |
Source and conversion
- Source checkpoint is FP8 e4m3 with block-128 scales (
weight_scale_inv); no BF16 of this finetune exists. Every tensor was dequantized block-wise on GPU to full precision before trellis encoding; scale tensors are consumed by the dequant and do not ship. FP8 noise (~2% per block) is second-order against the K2 trellis quantization error. - Encoded with exllamav3 v0.0.43
quantize_exl3(Hessian-free / uncalibrated,sigma_reg=0.025,mcgvariant), one tensor at a time, deterministic seeds (PYTHONHASHSEED=0, per-tensor seed from the tensor name). - ABI-compatible with the GLM-5.3-Flash EXL3 packs in the Mia overlay
convention: per-matrix payload byte-identical to the base-model encode
(2,109,444 B per expert matrix;
k_words=32;mcgstored as int32 shape(1,)=0xCBAC1FED). - Every shard verified post-encode: all expert keys present as
.trellis/.suh/.svh/.mcg, all native keys copied. - Original FP8-era
model.safetensors.index.jsonis preserved underreference/for provenance; it does not describe this pack's tensors.
How to run
This is not loadable by stock exllamav3/TabbyAPI (upstream has no
glm5_next architecture class yet) nor by llama.cpp (GGUF-only). It targets
the vLLM EXL3 overlay used by the GLM-5.3-Flash EXL3 community packs
(fused MoE kernels dispatch K per layer from tensor shapes at load). A
~91 GiB weight footprint fits a single 121 GiB unified-memory device
(e.g. DGX Spark / GB10) with ~30 GiB left for KV cache and runtime.
Related repositories
- vcruz305/GLM-5.3-Flash-Uncensored-EXL3-MixedK — mixed-precision variant: the most quantization-sensitive MoE layers re-encoded at K=3, selected by a per-layer K2-vs-K3 sensitivity scan.
- vcruz305/glm53-uncensored-exl3-k3-delta — the K3 layer overlay, displaced K2 originals, and the sensitivity evidence.
- vcruz305/GLM-5.3-Flash-EXL3-K2 — same recipe applied to the non-abliterated base model.
Limitations
- Not yet boot-tested. Structural and ABI verification passed; an
end-to-end
/v1load + generation check on the overlay runtime is the remaining qualification step. Treat as experimental until then. - K=2 is an aggressive rate. Expect measurable quality loss vs the FP8 source; no imatrix/Hessian calibration was used. Task benchmarks pending.
- Double quantization: FP8 → trellis. Unavoidable (no BF16 source exists); believed minor at this rate, not yet separately measured.
- Vision tower ships in BF16 but multimodal serving depends on the runtime; text-only serving is the validated intent.
- Uncensored model: safety behaviors of the base model were deliberately removed by the upstream finetune. You are responsible for your deployment and jurisdiction. Provided under MIT per the upstream license.