keys-Qwen3.8 27B NVFP4 DFlash2 Ablit+Cybersecurity Unlock 1M Context YaRN Single-DGXSpark
Best of the community, open source — compiled in one place.
This is not a new model Keys trained. It is a measured one-DGX-Spark recipe that wires together the best public pieces for Qwen3.8-27B: AEON-7’s uncensored abliteration, a uniform NVFP4 that actually fits GB10, Inco/Z Lab DFlash2 n=7, AEON’s prebuilt vLLM image, and YaRN to 1,048,576 tokens.
Go read the original repos. That is where the research, ablit trials, drafter, and Spark image live. We only put the wiring and the bake-off numbers in one folder.
| Hardware | one NVIDIA DGX Spark (GB10) — not the dual-Spark Flash-Next pair |
| Window | 1,048,576 (YaRN factor 4.0 × native 262,144) |
| 1M needle | HIT at 999,714 prompt tokens |
| Ablit + cyber | AEON uncensored weights + Keys unlock template → cyber 8/8, refusal32 32/32 |
| Speed (16k / seqs 64) | tea 135.7 agg @ c=16 · prose decode ~20 · code ~42 |
| Intelligence | model corrected a wrong harness key (Monday, not the gold Sunday) |
Prebuilt image (this is the upload):
ghcr.io/drowzeys/keys-qwen38-27b-nvfp4-dflash2-ablit-cyber-unlock-1m-yarn-single-dgxspark:latest
# digest sha256:fd31bd450510bebbb6d28623eb88ff8a86a6ac6c8db2a9009f21fafe8ff1eff9
That image is a retag of AEON-7’s ghcr.io/aeon-7/aeon-vllm-ultimate:latest
upstream digest sha256:dd2018473ed88bc23b01cfc3179b5b6896a7f0f152ae06d8d274db62d330ef48
(v0.27.1+aeon.sm121a.dspark). We did not rebuild vLLM. Layers are mounted from AEON-7.
Weights are not in this Hugging Face repo. Download them from the people who made them:
hf download sakamakismile/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4 \
--local-dir ~/models/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4
hf download incoai/Qwen3.8-27B-DFlash2 \
--local-dir ~/models/Qwen3.8-27B-DFlash2
docker pull ghcr.io/drowzeys/keys-qwen38-27b-nvfp4-dflash2-ablit-cyber-unlock-1m-yarn-single-dgxspark:latest
# identical to:
# docker pull ghcr.io/aeon-7/aeon-vllm-ultimate:latest
Why each community repo (read them)
Full table: CREDITS.md. Short version:
| We used | Because it is the best open piece for this job | Link |
|---|---|---|
| Qwen/Qwen3.8-27B | Best open dense 27B: code, agents, 262k rope, vision, thinking | HF |
| AEON-7 Ultimate Uncensored BF16 | Best coherence-first ablit of that 27B (not a lobotomy). This is the abliteration. | HF · GitHub |
| sakamakismile NVFP4 | Best uniform compressed-tensors NVFP4 of the AEON master that fits one Spark with DFlash2 + 1M KV | HF |
| AEON-7 NVFP4-MIXED | Best official Spark/5090 AEON quant — still go there. We used uniform on this box because it was +33% tea vs MIXED @ c=16 | HF |
| Inco + Z Lab DFlash2 | Best lossless block-diffusion drafter for Qwen3.8-27B; AEON Spark recipe is n=7 | incoai · z-lab · blog · code |
| AEON aeon-vllm-ultimate | Best prebuilt GB10 vLLM (TRITON_ATTN, DFlash, nested YaRN, qwen3 parsers) | ghcr.io/aeon-7/aeon-vllm-ultimate:latest |
| vLLM | Best production OpenAI-compatible server this can sit on | github.com/vllm-project/vllm |
| YaRN | Best standard RoPE stretch: 262144 × 4 = 1048576 | arXiv:2309.00071 |
Abliteration tools AEON used (read their card): abliterix, heretic, Arditi et al. 2406.11717, FernflowerAI SSM repair.
What Keys put together
- Nested YaRN under
text_config.rope_parametersand a YaRN overlay on the DFlash2config.json(vLLM will not stretch dflash rope past 262k by itself). - Cybersecurity-unlock default system template. AEON uncensored weights alone were 7/8 cyber / 30/32 refusal on our first B5 gate; after the template: 8/8 and 32/32.
- MIXED-style sampling (
repetition_penalty=1.05). - Measured eval on spark-13b3, 2026-09-08.
GMU 0.70 on the 1M seat, 0.75 on the 16k speed seat. Never above 0.85.
One-shot (one DGX Spark)
Accept RESPONSIBLE_USE.md, then:
git clone --depth 1 https://github.com/drowzeys/keys-Qwen3.8-27B-NVFP4-DFlash2-Ablit-Cybersecurity-Unlock-1M-Context-YARN-Single-DGXSpark
cd keys-Qwen3.8-27B-NVFP4-DFlash2-Ablit-Cybersecurity-Unlock-1M-Context-YARN-Single-DGXSpark
I_AGREE=1 bash one-shot.sh
That script: prints credits → pulls the AEON-7 prebuild (Keys GHCR, fallback to public aeon-vllm-ultimate) → downloads sakamakismile NVFP4 + Inco DFlash2 → launches 1M YaRN / seqs 2 / gmu 0.70 / DFlash2 n=7 → waits for max_model_len=1048576 → smokes PONG.
I_AGREE=1 bash one-shot.sh --speed # 16k / seqs 64 speed seat
I_AGREE=1 bash one-shot.sh --skip-download # weights already on disk
Health: curl -s http://127.0.0.1:8000/v1/models → aeon, max_model_len 1048576.
Short-prompt decode is not killed by the 1M cap; concurrency is (seqs 2 vs 64). Do not quote 135.7 agg on the 1M boot. GMU never above 0.85.
Measured (spark-13b3, 2026-09-08)
Engine vLLM 0.27.1+aeon.sm121a.dspark. Thinking off on eval.
16k / seqs 64 / gmu 0.75 (speed seat):
| Tea wall | 17.9 tok/s |
| Prose decode | ~19–21 tok/s |
| Code decode | 42.3 tok/s |
| Tea c=1 / 8 / 16 | 24.3 / 99.1 / 135.7 agg |
| Cat-6 mixed @ c=16 | 126.3 agg |
Capability: math 8/8 · coding 5/5 actual (whitespace harness miss) · intelligence 6/6 actual · truth 8/8 · agentic 4/4 · cyber 8/8 · refusal32 32/32.
The intelligence monday item: harness gold was Sunday. The model answered Monday. Yesterday Friday → today Saturday → tomorrow Sunday → day after tomorrow Monday. That is the model correcting a wrong exam key, not a miss. MIXED “passed” by matching the bad gold.
vs AEON Ultimate MIXED (2026-09-09)
Same-day bake-off: this uniform NVFP4 B5 (.4, 16k/seqs 64) vs AEON-7 NVFP4-MIXED (.1, 16k/seqs 32). Full tables: COMPARE.md.
| Keys B5 | AEON MIXED | |
|---|---|---|
| STEM 50-cat (math/physics/chem/genomics/folding/compare) | 39/40 | 39/40 |
| Protein folding | 5/5 | 5/5 |
monday actual |
Monday | Sunday |
| Cyber 10 + refusal32 | 10/10 · 32/32 | 10/10 · 32/32 |
| Essay / list / tea decode tok/s | 21.5 / 19.7 / 22.2 | 18.3 / 17.3 / 16.0 |
| Tea agg c=16 / c=32 | 147.7 / 181.8 | 117.4 / 152.7 |
Tied on STEM and unlock. B5 wins prose speed and the concurrency sweep. MIXED got the KE item (16 J); B5 dropped the ½ (32). MIXED still matches the bad Sunday gold.
1M YaRN / seqs 2 / gmu 0.70, KV 1,113,770 tokens:
| Prompt tok | Needle |
|---|---|
| 23,546 … 279,773 | HIT (280k is already past native 262,144) |
| 999,714 | HIT NX-7459C85377 (2.89 h prefill on one Spark) |
Full tables: RESULTS.md. No payloads in the public pack.
⚠️ Responsible use
Uncensored + cyber unlock. You own prompts, outputs, and harm. Gated access. See RESPONSIBLE_USE.md.
Files
| File | What |
|---|---|
COMPARE.md |
Keys B5 vs AEON Ultimate MIXED (2026-09-09 bake-off) |
one-shot.sh |
pull image + original weights + launch + smoke |
serve_b5_on_4.sh |
1M YaRN live recipe |
serve_b5_16k.sh |
16k / seqs 64 rollback |
yarn_hf_overrides.json |
nested YaRN hf-overrides |
dflash2_config_yarn_1m.json |
drafter rope overlay |
chat_template_b5_uncensored.jinja |
cyber unlock default system |
Dockerfile.prebuild |
FROM AEON digest (no rebuild) |
CREDITS.md |
every repo and why |
raw/ |
measured JSON |
License: Apache-2.0 inherited from Qwen. Image: AEON-7. Weights: AEON-7 / sakamakismile / Inco / Z Lab.