vcruz305/GLM-5.3-Flash-Uncensored-EXL3-K2

🤗 Hugging Face 来源text-generationmit98 GBother✓ 62 个校验和今天更新
需要做种者 →

GLM-5.3-Flash-Uncensored — EXL3 K2 (~2.0 bpw experts)

Sub-4-bit EXL3 (QTIP trellis) quantization of orcarouter/GLM-5.3-Flash-Uncensored-FP8, an abliterated finetune of zai-org/GLM-5.3-Flash (glm5_next / Glm5NextForConditionalGeneration). To our knowledge the first sub-4-bit EXL3 release of any uncensored GLM-5.3-Flash variant.

Sized for single-device serving: all 288 routed experts per MoE layer kept (no pruning), vision tower and NextN/MTP head kept, everything non-expert in native BF16.

Slab Format Approx. size
Routed experts (43 MoE layers × 288 experts × 3 projections = 37,152 matrices) EXL3 trellis, K=2 (~2.01 bpw payload) ~73 GiB
Attention / KDA / shared experts / routers / embeddings / lm_head / vision / MTP BF16 (native) ~18 GiB
Total weights ~91 GiB

Source and conversion

  • Source checkpoint is FP8 e4m3 with block-128 scales (weight_scale_inv); no BF16 of this finetune exists. Every tensor was dequantized block-wise on GPU to full precision before trellis encoding; scale tensors are consumed by the dequant and do not ship. FP8 noise (~2% per block) is second-order against the K2 trellis quantization error.
  • Encoded with exllamav3 v0.0.43 quantize_exl3 (Hessian-free / uncalibrated, sigma_reg=0.025, mcg variant), one tensor at a time, deterministic seeds (PYTHONHASHSEED=0, per-tensor seed from the tensor name).
  • ABI-compatible with the GLM-5.3-Flash EXL3 packs in the Mia overlay convention: per-matrix payload byte-identical to the base-model encode (2,109,444 B per expert matrix; k_words=32; mcg stored as int32 shape (1,) = 0xCBAC1FED).
  • Every shard verified post-encode: all expert keys present as .trellis/.suh/.svh/.mcg, all native keys copied.
  • Original FP8-era model.safetensors.index.json is preserved under reference/ for provenance; it does not describe this pack's tensors.

How to run

This is not loadable by stock exllamav3/TabbyAPI (upstream has no glm5_next architecture class yet) nor by llama.cpp (GGUF-only). It targets the vLLM EXL3 overlay used by the GLM-5.3-Flash EXL3 community packs (fused MoE kernels dispatch K per layer from tensor shapes at load). A ~91 GiB weight footprint fits a single 121 GiB unified-memory device (e.g. DGX Spark / GB10) with ~30 GiB left for KV cache and runtime.

Related repositories

Limitations

  1. Not yet boot-tested. Structural and ABI verification passed; an end-to-end /v1 load + generation check on the overlay runtime is the remaining qualification step. Treat as experimental until then.
  2. K=2 is an aggressive rate. Expect measurable quality loss vs the FP8 source; no imatrix/Hessian calibration was used. Task benchmarks pending.
  3. Double quantization: FP8 → trellis. Unavoidable (no BF16 source exists); believed minor at this rate, not yet separately measured.
  4. Vision tower ships in BF16 but multimodal serving depends on the runtime; text-only serving is the validated intent.
  5. Uncensored model: safety behaviors of the base model were deliberately removed by the upstream finetune. You are responsible for your deployment and jurisdiction. Provided under MIT per the upstream license.