vcruz305/Qwen3.8-27B-DFlash2-SAGE-EXL3-5.0bpw

🤗 Hugging Face 来源text-generationapache-2.0734M 参数1.5 GBsafetensors✓ 1 个校验和今天更新
需要做种者 →

Qwen3.8-27B-DFlash2 · SAGE-EXL3 · ~5.0 bpw

Made with SAGE, our own mixed-precision quantization method for EXL3 (not stock uniform convert.py). Held-out fidelity versus the BF16 draft is unscored. I am not claiming universal quality superiority or unbeaten KLD.

This repo is only the ~1.47 GB SAGE EXL3 draft of z-lab/Qwen3.8-27B-DFlash2. It is not a 27B. It is not AEON. Pair it with an EXL3 Qwen3.8 / AEON target. Stock ExLlamaV3 does not implement DFlash2.

AEON's Spark seat uses DFlash2 n=7 on the uncensored Qwen3.8-27B family: willfully compliant, still a real model when answers get long. RTX uses MTP n=3. Never both. This draft is for an EXL3 target, not for NVFP4/vLLM.

What it is

  • Format: EXL3 (quant_method: exl3)
  • Target budget: 5.0 bpw
  • Realized: 4.996 bpw mixed-K (4×K4, 12×K5, 9×K6, 9×K7, 3×K8)
  • Hub quantization_config.bits is the integer 5 because Hugging Face requires bits to be an int. The true rate is bits_per_weight: 4.996. Do not treat bits as the measured bpw.
  • Serialized: model.safetensors 1,468,305,753 bytes (~1.47 GB)
  • 5-layer DFlash2 drafter. Shares the target tokenizer / embeddings / lm_head.

What it is not

  • Not a standalone chat model
  • Not a target. Do not load it as the 27B.
  • Not NVFP4
  • Not a held-out quality score

Pairing

Recommended target:

vcruz305/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-sage-exl3-5.5bpw

Load with ExLlamaV3, not Transformers. You need a runtime that actually implements DFlash2.

On one GB10, greedy, ndt=7, fp16 KV, cs=2048: 26.15 tok/s easy continuation. Thinking prompts were slower, about 16 tok/s. ndt=15 crashed. That is a serving-stack receipt on that box, not a claim that EXL3 beats NVFP4.

Files

  • model.safetensors
  • config.json
  • quantization_config.json

Tokenizer files are omitted on purpose. Load the target tokenizer.

License and use

Apache-2.0, same as the upstream DFlash2 draft and Qwen/Qwen3.8-27B.

When you pair this with an AEON EXL3 target, you inherit that target's uncensored behavior. You own the prompt and the output. See the AEON BF16 and NVFP4-MIXED cards for the full user-responsibility text.

Credit: z-lab for DFlash2, AEON-7 and the Qwen authors for the target lineage. SAGE is mine.