Qwen3.8-27B AEON Ultimate Uncensored · SAGE-EXL3 5.5bpw
This is SAGE, my custom internal EXL3 path, not stock uniform ExLlama convert.py. I measure live activations (NativeReplay, mixed-K allocation), then assign a different K per linear so the budget goes where it matters. The goal is the highest-quality EXL3 I can get at this bpw, not a blunt equal-bit dump. I am not claiming universal quality superiority or unbeaten KLD.
This repo is my SAGE quant of AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16. I did not train a new 27B. I did not dequant-edit-requant an NVFP4 lattice.
AEON's intent, which this cut is meant to keep: uncensored for coherence and better answers, not a vanity KL of zero. Abliteration that does not lobotomize. Vision and the native MTP head left as the base. That BF16 is the master a deploy cut is baked from.
The NVFP4-MIXED sibling is AEON's willfully compliant deploy knife: still a real model when answers get long, mixed lattice not blunt NVFP4, seats on Spark / 5090 / 6000. Spark speculative is DFlash2 n=7. RTX is MTP n=3. Never both.
Model family
| Repo | Size | What it is |
|---|---|---|
| AEON BF16 master | ~54G | Full-precision uncensored teacher. Coherence over a vanity KL of zero. Vision + MTP untouched. |
| AEON NVFP4-MIXED | ~23.8G | Mixed NVFP4+FP8+BF16 deploy cut. vLLM on Spark / 5090 / 6000. DFlash2 n=7 on Spark, MTP n=3 on RTX. Never both. |
| This SAGE-EXL3 | 23.65 GB (3 shards) | My SAGE mixed-K EXL3 of the BF16 master. ExLlamaV3 only. |
What's in the weights
- Format: native ExLlamaV3 EXL3 (
quant_method: exl3) - Body rate: 5.5 bpw in
quantization_config.bits_per_weight - Hub
quantization_config.bitsis the integer 5 because Hugging Face requiresbitsto be an int. That field is not the true rate. Loaders should readbits_per_weight. - Embeddings, lm_head, MTP, and vision stay 16-bit
- 3 safetensors shards, 23.65 GB
- Tokenizer, chat template, and processor files come from the source. Vision is not runtime-tested here.
Load this with ExLlamaV3, not a plain Transformers weight loader.
Pairing
I run this as the EXL3 target with a DFlash2 draft. Spark-style speculative is DFlash n=7. Do not stack MTP and DFlash.
Matching SAGE draft: vcruz305/Qwen3.8-27B-DFlash2-SAGE-EXL3-5.0bpw
Stock ExLlamaV3 does not implement DFlash2. You need a runtime that does.
On one GB10, greedy, ndt=7, fp16 KV, cs=2048: 26.15 tok/s easy continuation. Thinking prompts were slower, about 16 tok/s. ndt=15 crashed. That is a serving-stack receipt on that box, not a claim that EXL3 beats NVFP4.
Source-relative KL (quant vs AEON, not smash vs Qwen)
AEON's smash KL (~0.099 nats/token on 100 harmless prompts, first 3 teacher-forced tokens) is abliteration vs stock Qwen. That number lives on his BF16 card. It is not a quant score.
This is a different protocol: did SAGE-EXL3 keep his BF16, not did abliteration move off Qwen.
teacher ExLlama FP16 operational view of AEON-7 BF16 (not native HF BF16)
student this SAGE-EXL3 5.5bpw
data 8 held-out SFT packs x 256 tokens
from resume-20260908/test_sft.jsonl
packed separately from calibration; first token of each pack not scored
positions 2040 next-token
metric mean token-level KL, student vs teacher next-token distribution
result mean KL 0.0033932283406782563
student NLL 2.719902172215992 (student-only; teacher NLL not published)
not WikiText, Sixcat, NVFP4, DFlash2, or AEON smash KL
0.00339 is tiny source-relative drift (quant tracking the teacher). It is not "better smash KL than 0.099." Do not mix the two.
License and use
Apache-2.0, inherited from Qwen/Qwen3.8-27B.
This inherits AEON's uncensored behavior. Outputs can be wrong, biased, or disallowed where you live. You own the prompt, the filter, and the fallout. AEON's full user-responsibility text lives on the BF16 and NVFP4-MIXED cards. Read those. I am not copying it here.
Credit: AEON-7 and the Qwen authors for the model. ExLlamaV3 for the format. SAGE is mine.