msuiche/Qwen3.8-Flash-Next-abliterated-cyber-GLP-47

🤗 Hugging Face 来源mit8 MBGGUF✓ 1 个校验和今天更新
帮助这个模型通过 Pirate Face 分发

模型卡、文件列表和校验和已在此收录。如果你持有文件并有权分享,可以提交种子,让其他人从节点下载。

获取来源文件需要 Hugging Face 批准。
帮助为此模型做种

Qwen3.8-Flash-Next — abliterated refusal direction (GLP-47)

A GLP control vector (GGUF Layer Projection, glp.mode=project) for Qwen/Qwen3.8-Flash-Next. 47 per-layer unit directions over the widened hyper-connection residual stream (10240 = 4 x 2560), layers 1-47, fp32. No model weights are modified; a conforming reader applies h <- h - alpha * (h . d) d at the post-layer residual stream.

Derived day-0 (2026-08-26) with captain-vector from a public contrast (AdvBench harmful_behaviors vs Stanford Alpaca), difference-of-means, per-layer structure, on the BF16 reference checkpoint (revision f5d08274bafd880402bd16f5e3e6c514136ec06c). Engine-independence verified: a vLLM-lane capture reproduces this direction at cosine 0.9931.

Measured effect (greedy, 400 tokens, n=32 per suite)

alpha refusal32 delivered cyber32 delivered benign32 capability12
0.0 (stock) 1/32 (3.1%) 5/32 (15.6%) 31/32 12/12
1.0 (shipped default) 26/32 (81.2%) 32/32 (100%) 32/32 12/12
1.5 24/32 — 32/32 12/12
2.0 24/32 — 31/32 12/12

alpha=1.0 is the measured peak; higher alpha over-projects and refusal32 delivery DROPS. The cyber holdout saturates completely at alpha=1.0 — the remaining refusal32 holdouts are a physical-harm/petty-crime cluster, not cyber. This is v1, not a 32/32 refusal32 claim.

Usage

Confirmed bases

  • Derived from Qwen/Qwen3.8-Flash-Next BF16 reference (revision f5d08274bafd880402bd16f5e3e6c514136ec06c), AdvBench-32 vs Alpaca-32 contrast, difference-of-means per layer.
  • Validated serving: RadixArk/Qwen3.8-Flash-Next-NVFP4 on the day-0 vllm/vllm-openai:qwen38-flash-next image (B200): refusal32 25/32 steered vs 1/32 stock — quantization does not degrade the direction.

Option 1 — weightless wizard (recommended)

git clone https://github.com/msuiche/weightless.git && cd weightless && python3 setup.py
# pick: "Qwen3.8-Flash-Next TP=2 serving" (2x DGX Spark lane)

The wizard downloads this vector, structure-checks the hotfix, deploys over ssh, and smoke-tests the endpoint.

Option 2 — manual (any vLLM container)

huggingface-cli download msuiche/Qwen3.8-Flash-Next-abliterated-cyber-GLP-47 --include "*.gguf"
# inside the serving container, BEFORE vllm serve:
WEIGHTLESS_STEER_PATH=/cache/huggingface/Qwen3.8-Flash-Next-abliterated-cyber-GLP-47-L1-47-a1.gguf \
  python3 /patches/hotfix-qwen38fn-steering-projective.py && exec vllm serve ...

The hotfix (in weightless/patches/) is fail-closed: if its anchors don't match the vLLM build it aborts before serving rather than running unsteered. On the RadixArk NVFP4 checkpoint the day-0 image also needs patches/patch-qwen38fn-ple-fp8-nvfp4.py (the PLE n-gram table is FP8-serialized) — the wizard applies both.

Option 3 — other runtimes

Spec-conformant GGUF control vector (glp.mode=project, spec: spec/GLP.md in the weightless repo). Stock llama.cpp's control-vector apply is additive, not projective — it is NOT a conforming reader for this file.

alpha is a runtime parameter, never folded into the vector; 1.0 is the measured peak — higher over-projects.

Content SHA-256 (tensor bytes): 116e3a6cd4645255f1469f53b86050cebf922a22db90ec88ed00b63d24468fb9