msuiche/Qwen3.8-27B-hedging-GLP-63-L1-63-a0.5

🤗 Hugging Face 来源mit激活 27B10 MBGGUF✓ 2 个校验和今天更新
帮助这个模型通过 Pirate Face 分发

模型卡、文件列表和校验和已在此收录。如果你持有文件并有权分享,可以提交种子,让其他人从节点下载。

获取来源文件需要 Hugging Face 批准。
帮助为此模型做种

Qwen3.8-27B hedging GLP-63 — the second direction (companion to GLP-49)

A 1.3 MB control vector that targets the residual soft rejections — disclaimers and warnings attached to complying answers, seed-dependent in the chain-of-thought — that remain after msuiche/Qwen3.8-27B-abliterated-cyber-GLP-49 is applied. GLP-49 removes the dominant refusal direction (the trained category gate). This is the next component of the subspace: the near-threshold hedging that sampling flips on and off.

Requested and observed in discussion #1 on the GLP-49 repo.

file size needs
Qwen3.8-27B-hedging-GLP-63-L1-63-a0.5.gguf 1.3 MB a runtime that understands glp.mode = project
adapter_model.safetensors + adapter_config.json 8.6 MB stock peft — see "PEFT LoRA export" below

The base checkpoint is untouched. No weights are redistributed.


Results

Measured on Qwen/Qwen3.8-27B at revision 1d4bf0f2ff60, GLP-49 armed at α=1.0 in every arm (this vector is derived and validated on the GLP-49-steered model — that is the state where the residual hedging lives), thinking enabled, temp 1.0 / top_p 0.95 / top_k 20 (the checkpoint's own generation defaults), 69 hedge-prone prompts × 8 seeds, paired across arms. "CoT-marker" = a speaker-stance disclaimer phrase ("I must warn", "for educational purposes", "please be aware", …) in the reasoning trace of a sample that produced a substantive answer.

α₂ (this vector) CoT-marker rate all-suite hedge share benign32-holdout cyber32 delivered collapse
0 (GLP-49 only) 0.372 0.410 29/32 clean 16/32 0/552
0.5 (shipped default) 0.258 0.366 30–31/32 clean 16/32 0/552
1.0 0.244 0.333 31/32 clean 14/32 5/552
2.0 0.208 0.306 30/32 clean 21/32 9/552

On the strictly seed-dependent subpopulation (46 prompts that hedge on some seeds and not others — the phenomenon reported in the discussion): CoT-marker 0.290 → 0.208 at α=0.5.

Also measured at α=0.5: argumentative premise-rejections (the "I cannot provide instructions on how to make sourdough rise improperly" state the GLP-49 card documents) drop from 164 to 92 of 552 samples, and clean-stop rate rises from 81 % to 90 %.

This is a reduction, not an elimination. About a third of the CoT marker rate is gone at the shipped dose; the rest sits in components this one direction does not capture, and on deflection-type prompts the caution register survives in rephrased openings. Hedging decomposes as a subspace; this is the first component of it we could validate, not the whole of it.

Do not push α at this vector's hardest prompts

At α ≥ 1.0, on the most refusal-charged prompts (explosives, self-harm), the model emits Hmm, and stops — 5/552 samples at α=1.0, 9/552 at α=2.0, none at α=0.5. With both the refusal gate and the hedging component removed at full strength, the model has no engagement pathway left on those prompts and terminates instead of answering or refusing. That onset is why alpha_default is 0.5 despite α=1.0 being the exact removal point.


Usage

Apply after GLP-49, at the post-layer residual stream (hidden_states + residual), per layer:

h ← h − α₂ (h · d̂) d̂         α₂ = 0.5

The vector declares glp.mode = project. llama.cpp's built-in control vectors are additive (h ← h + s·d̂), a different operation that fails silently; a reader that does not understand glp.mode must refuse the file.

Combined serving needs two independent projections. The reference single-file hotfix (weightless hotfix-qwen38-steering-projective.py) applies one GGUF at one α, and its rank-k loader orthonormalises rows, so a merged two-row file would force GLP-49's α=1.0 and this vector's α=0.5 to share a dose. The dual-path hotfix used for every measurement above (GLP-49 slot 1 at α=1.0, this vector slot 2 at α₂) is published with the derivation record: refusal-research/experiments/20260910-qwen38-27b-hedging-dir/staging/patch_qwen38_hedge.py.

Checkpoint-bound. Valid for Qwen/Qwen3.8-27B at revision 1d4bf0f2ff60 and its re-quantisations. Derived on the GLP-49-steered model; applying it without GLP-49 is unmeasured.


PEFT LoRA export

adapter_model.safetensors + adapter_config.json are the rank-1 LoRA form of the same intervention, baked with the exact conventions of the GLP-49 adapter so the two merge identically and stack: for each residual-writing matmul (self_attn.o_proj on the 16 full-attention layers, linear_attn.out_proj on the other 47, mlp.down_proj on every layer, span L1-63),

lora_B = d̂                  (5120×1)
lora_A = −α₂ · d̂ᵀW          (1×n_in), α₂ = 0.5 baked in

lora_alpha / r = 1.0, so the dose lives entirely in lora_A — do not scale the adapter, and do not raise α₂ (the α ≥ 1.0 collapse described above applies here too). Apply after the GLP-49 adapter, or merge both into the same base in either order — the edits are disjoint rank-1 terms on the same matrices. merge_and_unload() on the stack, then GGUF conversion, works with stock tooling. Same checkpoint-binding as the GGUF: lora_A carries W of revision 1d4bf0f2ff60.

Two things to know before you read anything into baked results. The adapter is a troubleshooting/interop form — use it to check that the direction lands, or to probe a merge for survival; the GGUF plus a runtime hook is how the steering is actually served. And it is not bit-identical to that runtime path: the LoRA projects each residual writer's output, which by linearity equals projecting the layer's new contributions at α₂, but a d̂-component already sitting in the incoming stream (embeddings, any writer not covered) passes through, where the runtime hook removes it regardless of origin. Close in practice when writers re-inject the direction every layer — which the per-layer structure here does — but not the same object.


What it does to ordinary answers, measured honestly

  • Benign answers lose ~20 % of their length. Part of that is the caution scaffolding this vector exists to remove (a benign papermaking answer stopped opening with "I'd recommend not trying this at home" and now opens with the process); part is genuine elaboration (a cast-iron answer reads thinner). Read samples before shipping user-facing output.
  • Delivery of substantive technical answers rises on cyber32 (long-answer count 45 → 69 of 112 samples at α=1.0) — the model is not terser because it is collapsing; it is terser where it used to pad with disclaimers.
  • No new refusals on the benign holdout at any α tested.

Derivation (provenance)

base            Qwen/Qwen3.8-27B @ 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
derived on      the GLP-49-steered model (GLP-49 L10-58, α=1.0, armed)
contrast        clean-compliance vs disclaimer-laden-compliance,
                prompt-driven: classes from the steered model's own seeded
                generations (hedge lexicon + hand audit); within-prompt
                natural branch points (the model's own CoT prefix cut
                just before its first disclaimer vs a length-matched
                clean prefix of the same prompt). No forced/pinned text.
estimator       difference of means, per layer, 0.5% massive-activation
                masked (this checkpoint's dim-3994 pathology)
hook            residual_stream_post_layer (hidden_states + residual),
                post-steering tap; layers 1-63; derived_at == hook_point
gates           null-ratio ≥5x on 63/63 layers (top L52-63 at 12-13x),
                adjacent-layer cosine 0.94, |cos(d, GLP-49)| median 0.05
                (a new axis, not refusal residue)
content_sha256  4a4a3c3c850c446a6bd0b150a1879fa09f77ac9ae20bef37f0dd90750bf7912f
experiment      refusal-research/experiments/20260910-qwen38-27b-hedging-dir

Note on the design that did not work: the same contrast measured at the last prompt token (the "poised" capture that works for the refusal direction) separates hedgy from clean prompts at under 3× the shuffled-label null — below the 5× ship gate, at every layer. The hedging propensity is not committed at prefill; it crystallises during the CoT. That is consistent with the seed-dependence that motivated this vector, and it is why the shipped direction is derived at natural branch points instead.

Intended use

Security research and evaluation. Measuring how robustly a behaviour is gated requires being able to switch the gate's components off one at a time and observe what changes, and that work is only possible on open-weight models.