Qwen3.8-27B-DFlash2 · SAGE-EXL3 · ~5.0 bpw
Made with SAGE, our own mixed-precision quantization method for EXL3 (not stock uniform convert.py). Held-out fidelity versus the BF16 draft is unscored. I am not claiming universal quality superiority or unbeaten KLD.
This repo is only the ~1.47 GB SAGE EXL3 draft of z-lab/Qwen3.8-27B-DFlash2. It is not a 27B. It is not AEON. Pair it with an EXL3 Qwen3.8 / AEON target. Stock ExLlamaV3 does not implement DFlash2.
AEON's Spark seat uses DFlash2 n=7 on the uncensored Qwen3.8-27B family: willfully compliant, still a real model when answers get long. RTX uses MTP n=3. Never both. This draft is for an EXL3 target, not for NVFP4/vLLM.
What it is
- Format: EXL3 (
quant_method: exl3) - Target budget: 5.0 bpw
- Realized: 4.996 bpw mixed-K (
4×K4, 12×K5, 9×K6, 9×K7, 3×K8) - Hub
quantization_config.bitsis the integer 5 because Hugging Face requiresbitsto be an int. The true rate isbits_per_weight: 4.996. Do not treatbitsas the measured bpw. - Serialized:
model.safetensors1,468,305,753 bytes (~1.47 GB) - 5-layer DFlash2 drafter. Shares the target tokenizer / embeddings / lm_head.
What it is not
- Not a standalone chat model
- Not a target. Do not load it as the 27B.
- Not NVFP4
- Not a held-out quality score
Pairing
Recommended target:
vcruz305/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-sage-exl3-5.5bpw
Load with ExLlamaV3, not Transformers. You need a runtime that actually implements DFlash2.
On one GB10, greedy, ndt=7, fp16 KV, cs=2048: 26.15 tok/s easy continuation. Thinking prompts were slower, about 16 tok/s. ndt=15 crashed. That is a serving-stack receipt on that box, not a claim that EXL3 beats NVFP4.
Files
model.safetensorsconfig.jsonquantization_config.json
Tokenizer files are omitted on purpose. Load the target tokenizer.
License and use
Apache-2.0, same as the upstream DFlash2 draft and Qwen/Qwen3.8-27B.
When you pair this with an AEON EXL3 target, you inherit that target's uncensored behavior. You own the prompt and the output. See the AEON BF16 and NVFP4-MIXED cards for the full user-responsibility text.
Credit: z-lab for DFlash2, AEON-7 and the Qwen authors for the target lineage. SAGE is mine.