GLM-5.3-Flash EXL3 K2/K3 mix
A mixed-precision EXL3 pack of zai-org/GLM-5.3-Flash-BF16: the K2 base (vcruz305/GLM-5.3-Flash-EXL3-K2) with six routed-expert layers promoted to K3. Everything else is byte-identical to the base. Attention, shared experts, embeddings, head and vision stay source-native BF16.
| Base | vcruz305/GLM-5.3-Flash-EXL3-K2, 2-bit MCG trellis, routed experts only |
| K3 layers | 24, 27, 35, 37, 42, 45 (all 288 routed experts, all three projections each) |
| Routed-expert width | 37 layers x K2 + 6 layers x K3 = 2.14 bpw effective |
| Architecture | Glm5NextForConditionalGeneration |
| Shards | 120; 98 identical to the base, 22 rewritten |
| Tensors replaced | 20,736 (5,184 expert weights x 4 EXL3 tensors) |
| Target | one NVIDIA DGX Spark / GB10 (SM121), vLLM --quantization exl3 |
Composition
Six routed-expert layers (24, 27, 35, 37, 42, 45) are K3; the other 37 are the
K2 base. Every promoted layer keeps gate_proj, up_proj and down_proj at
the same K, which the fused MoE kernel requires. Unchanged shards are
byte-identical to the base.
quantization_config keeps bits=2 as the baseline and adds:
"layer_bits": {"24": 3, "27": 3, "35": 3, "37": 3, "42": 3, "45": 3}
Serving: needs per-layer K support
The EXL3 plugin allocates trellis parameters from bits, so a stock K2 plugin fails to load this pack with EXL3 load shape mismatch on the K3 layers. It fails loudly, not silently. The plugin in the recipe repo reads layer_bits and dispatches K per layer; the fused exl3_moe kernel takes K per launch, so K2 and K3 layers coexist in one fused path.
git clone https://github.com/vcruz305/GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe
cd GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe
python scripts/preflight.py
bash scripts/install_prebuilt.sh # runtime wheels from vcruz305/GLM-5.3-Flash-EXL3-K2-spark-vllm
hf download vcruz305/GLM-5.3-Flash-EXL3-K2K3-mix --local-dir ~/models/GLM-5.3-Flash-EXL3-K2K3-mix
MODEL_DIR=~/models/GLM-5.3-Flash-EXL3-K2K3-mix SPEC_METHOD=mtp MTP_TOKENS=2 \
MAX_MODEL_LEN=65536 MAX_NUM_SEQS=1 bash scripts/serve_one_spark.sh
Load log must show EXL3 per-layer K override ... bits=3 for the six layers and fused_moe=exl3_moe.
Status
| Gate | Result |
|---|---|
| Shard validation | done: 22 shards carry K3 tensors, 98 identical to the base, K3 trellis shapes [256, 128, 48] confirmed |
| Load with per-layer K plugin | done: loads on the 2026-08-30 wheels, 0 shape mismatches |
| Memory at 64k | done: 98.02 GiB after load (K2: 92.31), KV cache 337,042 tokens at util 0.91 (K2: 786,432) |
| Fused K2/K3 dispatch | done: K3 layers 24/27/35/37/42/45 load into K3-allocated parameters and feed the fused kernel at k=3 |
| Speed vs K2 (same four-workload ladder, 64k) | done: 17.19 vs 16.74 tok/s (+2.7%); faster on code (+5.7%), math (+6.1%), structured (+1.6%), slower on prose (-2.5%); MTP acceptance up on every workload but prose |
| sixcat 0.5.1, think-on, 64k | done: overall 83.33 (knowledge 65, math 100, truth 85, instruct 75, code 85, tools 90), 120/120, no faults; K2 on the same runtime: 84.17, one item apart, which is this suite's noise floor. Receipt in the recipe's docs/SIXCAT.md |
| KLD vs BF16 teacher (fidelity suite v1) | 512 contexts, 1,048,064 positions, full-vocab KL through the shared head: mix 0.3121 nats (95% CI 0.299–0.325), top-1 0.795 vs K2 0.3346 / 0.788 on the same contexts. Paired: mix lower by 0.0225 nats (CI [0.0207, 0.0243]) and lower on 505 of 512 contexts. The suite's FP8 anchor is 0.028 / 0.943. Method and scorer validation in the recipe's docs/KLD.md |
Weights are uploaded once the load and dispatch gates pass. Until the table above is filled, treat this card as a build record, not a qualified release.
Known runtime issue
The vLLM glm5next runtime this pack is served on has an open out-of-bounds write in the sparse-MLA K-pool tail on hybrid models. It is not EXL3-specific and affects K2, this mix, and TR3 packs alike. Root cause, reproducer, detector and fix: docs/KPOOL_TAIL_BUG.md in the recipe repo.
Related
| Repo | Role |
|---|---|
| GLM-5.3-Flash-EXL3-K2 | the immutable K2 base |
| GLM-5.3-Flash-EXL3-K2-spark-vllm | prebuilt runtime wheels for GB10 |
| GLM-5.3-Flash-EXL3-K2K3-mix-DGX-Spark-recipe | this pack's recipe: install, serve, bench, measurements vs K2 |
| GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe | the K2 base recipe |
License: MIT (Z.AI), same as the BF16 source. Community quant; not an official Z.ai release.