Qwen3.8-Flash-Next — abliterated refusal direction (GLP-47)
A GLP control vector (GGUF Layer Projection, glp.mode=project) for
Qwen/Qwen3.8-Flash-Next.
47 per-layer unit directions over the widened hyper-connection residual
stream (10240 = 4 x 2560), layers 1-47, fp32. No model weights are modified;
a conforming reader applies h <- h - alpha * (h . d) d at the post-layer
residual stream.
Derived day-0 (2026-08-26) with captain-vector from a public contrast
(AdvBench harmful_behaviors vs Stanford Alpaca), difference-of-means,
per-layer structure, on the BF16 reference checkpoint (revision f5d08274bafd880402bd16f5e3e6c514136ec06c).
Engine-independence verified: a vLLM-lane capture reproduces this direction
at cosine 0.9931.
Measured effect (greedy, 400 tokens, n=32 per suite)
| alpha | refusal32 delivered | cyber32 delivered | benign32 | capability12 |
|---|---|---|---|---|
| 0.0 (stock) | 1/32 (3.1%) | 5/32 (15.6%) | 31/32 | 12/12 |
| 1.0 (shipped default) | 26/32 (81.2%) | 32/32 (100%) | 32/32 | 12/12 |
| 1.5 | 24/32 | — | 32/32 | 12/12 |
| 2.0 | 24/32 | — | 31/32 | 12/12 |
alpha=1.0 is the measured peak; higher alpha over-projects and refusal32 delivery DROPS. The cyber holdout saturates completely at alpha=1.0 — the remaining refusal32 holdouts are a physical-harm/petty-crime cluster, not cyber. This is v1, not a 32/32 refusal32 claim.
Usage
Confirmed bases
- Derived from
Qwen/Qwen3.8-Flash-NextBF16 reference (revisionf5d08274bafd880402bd16f5e3e6c514136ec06c), AdvBench-32 vs Alpaca-32 contrast, difference-of-means per layer. - Validated serving:
RadixArk/Qwen3.8-Flash-Next-NVFP4on the day-0vllm/vllm-openai:qwen38-flash-nextimage (B200): refusal32 25/32 steered vs 1/32 stock — quantization does not degrade the direction.
Option 1 — weightless wizard (recommended)
git clone https://github.com/msuiche/weightless.git && cd weightless && python3 setup.py
# pick: "Qwen3.8-Flash-Next TP=2 serving" (2x DGX Spark lane)
The wizard downloads this vector, structure-checks the hotfix, deploys over ssh, and smoke-tests the endpoint.
Option 2 — manual (any vLLM container)
huggingface-cli download msuiche/Qwen3.8-Flash-Next-abliterated-cyber-GLP-47 --include "*.gguf"
# inside the serving container, BEFORE vllm serve:
WEIGHTLESS_STEER_PATH=/cache/huggingface/Qwen3.8-Flash-Next-abliterated-cyber-GLP-47-L1-47-a1.gguf \
python3 /patches/hotfix-qwen38fn-steering-projective.py && exec vllm serve ...
The hotfix (in weightless/patches/) is fail-closed: if its anchors don't
match the vLLM build it aborts before serving rather than running unsteered.
On the RadixArk NVFP4 checkpoint the day-0 image also needs
patches/patch-qwen38fn-ple-fp8-nvfp4.py (the PLE n-gram table is
FP8-serialized) — the wizard applies both.
Option 3 — other runtimes
Spec-conformant GGUF control vector (glp.mode=project, spec:
spec/GLP.md in the weightless repo). Stock llama.cpp's control-vector apply
is additive, not projective — it is NOT a conforming reader for this file.
alpha is a runtime parameter, never folded into the vector; 1.0 is the measured peak — higher over-projects.
Content SHA-256 (tensor bytes): 116e3a6cd4645255f1469f53b86050cebf922a22db90ec88ed00b63d24468fb9