vcruz305/GLM-5.3-Flash-EXL3-K2K3-mix

🤗 Hugging Face sourcetext-generationmit51.6B params103 GBsafetensors✓ 122 checksumsupdated today
Needs seeder →

GLM-5.3-Flash EXL3 K2/K3 mix

A mixed-precision EXL3 pack of zai-org/GLM-5.3-Flash-BF16: the K2 base (vcruz305/GLM-5.3-Flash-EXL3-K2) with six routed-expert layers promoted to K3. Everything else is byte-identical to the base. Attention, shared experts, embeddings, head and vision stay source-native BF16.

Base vcruz305/GLM-5.3-Flash-EXL3-K2, 2-bit MCG trellis, routed experts only
K3 layers 24, 27, 35, 37, 42, 45 (all 288 routed experts, all three projections each)
Routed-expert width 37 layers x K2 + 6 layers x K3 = 2.14 bpw effective
Architecture Glm5NextForConditionalGeneration
Shards 120; 98 identical to the base, 22 rewritten
Tensors replaced 20,736 (5,184 expert weights x 4 EXL3 tensors)
Target one NVIDIA DGX Spark / GB10 (SM121), vLLM --quantization exl3

Composition

Six routed-expert layers (24, 27, 35, 37, 42, 45) are K3; the other 37 are the K2 base. Every promoted layer keeps gate_proj, up_proj and down_proj at the same K, which the fused MoE kernel requires. Unchanged shards are byte-identical to the base.

quantization_config keeps bits=2 as the baseline and adds:

"layer_bits": {"24": 3, "27": 3, "35": 3, "37": 3, "42": 3, "45": 3}

Serving: needs per-layer K support

The EXL3 plugin allocates trellis parameters from bits, so a stock K2 plugin fails to load this pack with EXL3 load shape mismatch on the K3 layers. It fails loudly, not silently. The plugin in the recipe repo reads layer_bits and dispatches K per layer; the fused exl3_moe kernel takes K per launch, so K2 and K3 layers coexist in one fused path.

git clone https://github.com/vcruz305/GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe
cd GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe
python scripts/preflight.py
bash scripts/install_prebuilt.sh      # runtime wheels from vcruz305/GLM-5.3-Flash-EXL3-K2-spark-vllm
hf download vcruz305/GLM-5.3-Flash-EXL3-K2K3-mix --local-dir ~/models/GLM-5.3-Flash-EXL3-K2K3-mix
MODEL_DIR=~/models/GLM-5.3-Flash-EXL3-K2K3-mix SPEC_METHOD=mtp MTP_TOKENS=2 \
  MAX_MODEL_LEN=65536 MAX_NUM_SEQS=1 bash scripts/serve_one_spark.sh

Load log must show EXL3 per-layer K override ... bits=3 for the six layers and fused_moe=exl3_moe.

Status

Gate Result
Shard validation done: 22 shards carry K3 tensors, 98 identical to the base, K3 trellis shapes [256, 128, 48] confirmed
Load with per-layer K plugin done: loads on the 2026-08-30 wheels, 0 shape mismatches
Memory at 64k done: 98.02 GiB after load (K2: 92.31), KV cache 337,042 tokens at util 0.91 (K2: 786,432)
Fused K2/K3 dispatch done: K3 layers 24/27/35/37/42/45 load into K3-allocated parameters and feed the fused kernel at k=3
Speed vs K2 (same four-workload ladder, 64k) done: 17.19 vs 16.74 tok/s (+2.7%); faster on code (+5.7%), math (+6.1%), structured (+1.6%), slower on prose (-2.5%); MTP acceptance up on every workload but prose
sixcat 0.5.1, think-on, 64k done: overall 83.33 (knowledge 65, math 100, truth 85, instruct 75, code 85, tools 90), 120/120, no faults; K2 on the same runtime: 84.17, one item apart, which is this suite's noise floor. Receipt in the recipe's docs/SIXCAT.md
KLD vs BF16 teacher (fidelity suite v1) 512 contexts, 1,048,064 positions, full-vocab KL through the shared head: mix 0.3121 nats (95% CI 0.299–0.325), top-1 0.795 vs K2 0.3346 / 0.788 on the same contexts. Paired: mix lower by 0.0225 nats (CI [0.0207, 0.0243]) and lower on 505 of 512 contexts. The suite's FP8 anchor is 0.028 / 0.943. Method and scorer validation in the recipe's docs/KLD.md

Weights are uploaded once the load and dispatch gates pass. Until the table above is filled, treat this card as a build record, not a qualified release.

Known runtime issue

The vLLM glm5next runtime this pack is served on has an open out-of-bounds write in the sparse-MLA K-pool tail on hybrid models. It is not EXL3-specific and affects K2, this mix, and TR3 packs alike. Root cause, reproducer, detector and fix: docs/KPOOL_TAIL_BUG.md in the recipe repo.

Related

Repo Role
GLM-5.3-Flash-EXL3-K2 the immutable K2 base
GLM-5.3-Flash-EXL3-K2-spark-vllm prebuilt runtime wheels for GB10
GLM-5.3-Flash-EXL3-K2K3-mix-DGX-Spark-recipe this pack's recipe: install, serve, bench, measurements vs K2
GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe the K2 base recipe

License: MIT (Z.AI), same as the BF16 source. Community quant; not an official Z.ai release.