Qwen3.8-27B hedging GLP-63 — the second direction (companion to GLP-49)
A 1.3 MB control vector that targets the residual soft rejections —
disclaimers and warnings attached to complying answers, seed-dependent in
the chain-of-thought — that remain after
msuiche/Qwen3.8-27B-abliterated-cyber-GLP-49
is applied. GLP-49 removes the dominant refusal direction (the trained
category gate). This is the next component of the subspace: the
near-threshold hedging that sampling flips on and off.
Requested and observed in discussion #1 on the GLP-49 repo.
| file | size | needs |
|---|---|---|
Qwen3.8-27B-hedging-GLP-63-L1-63-a0.5.gguf |
1.3 MB | a runtime that understands glp.mode = project |
adapter_model.safetensors + adapter_config.json |
8.6 MB | stock peft — see "PEFT LoRA export" below |
The base checkpoint is untouched. No weights are redistributed.
Results
Measured on Qwen/Qwen3.8-27B at revision 1d4bf0f2ff60, GLP-49 armed
at α=1.0 in every arm (this vector is derived and validated on the
GLP-49-steered model — that is the state where the residual hedging
lives), thinking enabled, temp 1.0 / top_p 0.95 / top_k 20 (the
checkpoint's own generation defaults), 69 hedge-prone prompts × 8 seeds,
paired across arms. "CoT-marker" = a speaker-stance disclaimer phrase
("I must warn", "for educational purposes", "please be aware", …) in the
reasoning trace of a sample that produced a substantive answer.
| α₂ (this vector) | CoT-marker rate | all-suite hedge share | benign32-holdout | cyber32 delivered | collapse |
|---|---|---|---|---|---|
| 0 (GLP-49 only) | 0.372 | 0.410 | 29/32 clean | 16/32 | 0/552 |
| 0.5 (shipped default) | 0.258 | 0.366 | 30–31/32 clean | 16/32 | 0/552 |
| 1.0 | 0.244 | 0.333 | 31/32 clean | 14/32 | 5/552 |
| 2.0 | 0.208 | 0.306 | 30/32 clean | 21/32 | 9/552 |
On the strictly seed-dependent subpopulation (46 prompts that hedge on some seeds and not others — the phenomenon reported in the discussion): CoT-marker 0.290 → 0.208 at α=0.5.
Also measured at α=0.5: argumentative premise-rejections (the "I cannot provide instructions on how to make sourdough rise improperly" state the GLP-49 card documents) drop from 164 to 92 of 552 samples, and clean-stop rate rises from 81 % to 90 %.
This is a reduction, not an elimination. About a third of the CoT marker rate is gone at the shipped dose; the rest sits in components this one direction does not capture, and on deflection-type prompts the caution register survives in rephrased openings. Hedging decomposes as a subspace; this is the first component of it we could validate, not the whole of it.
Do not push α at this vector's hardest prompts
At α ≥ 1.0, on the most refusal-charged prompts (explosives, self-harm),
the model emits Hmm, and stops — 5/552 samples at α=1.0, 9/552 at
α=2.0, none at α=0.5. With both the refusal gate and the hedging
component removed at full strength, the model has no engagement pathway
left on those prompts and terminates instead of answering or refusing.
That onset is why alpha_default is 0.5 despite α=1.0 being the exact
removal point.
Usage
Apply after GLP-49, at the post-layer residual stream
(hidden_states + residual), per layer:
h ← h − α₂ (h · d̂) d̂ α₂ = 0.5
The vector declares glp.mode = project. llama.cpp's built-in control
vectors are additive (h ← h + s·d̂), a different operation that fails
silently; a reader that does not understand glp.mode must refuse the
file.
Combined serving needs two independent projections. The reference
single-file hotfix
(weightless hotfix-qwen38-steering-projective.py)
applies one GGUF at one α, and its rank-k loader orthonormalises rows, so
a merged two-row file would force GLP-49's α=1.0 and this vector's α=0.5
to share a dose. The dual-path hotfix used for every measurement above
(GLP-49 slot 1 at α=1.0, this vector slot 2 at α₂) is published with the
derivation record:
refusal-research/experiments/20260910-qwen38-27b-hedging-dir/staging/patch_qwen38_hedge.py.
Checkpoint-bound. Valid for Qwen/Qwen3.8-27B at revision
1d4bf0f2ff60 and its re-quantisations. Derived on the GLP-49-steered
model; applying it without GLP-49 is unmeasured.
PEFT LoRA export
adapter_model.safetensors + adapter_config.json are the rank-1 LoRA
form of the same intervention, baked with the exact conventions of the
GLP-49 adapter so the two merge identically and stack: for each
residual-writing matmul (self_attn.o_proj on the 16 full-attention
layers, linear_attn.out_proj on the other 47, mlp.down_proj on every
layer, span L1-63),
lora_B = d̂ (5120×1)
lora_A = −α₂ · d̂ᵀW (1×n_in), α₂ = 0.5 baked in
lora_alpha / r = 1.0, so the dose lives entirely in lora_A — do not
scale the adapter, and do not raise α₂ (the α ≥ 1.0 collapse described
above applies here too). Apply after the GLP-49 adapter, or merge
both into the same base in either order — the edits are disjoint rank-1
terms on the same matrices. merge_and_unload() on the stack, then
GGUF conversion, works with stock tooling. Same checkpoint-binding as
the GGUF: lora_A carries W of revision 1d4bf0f2ff60.
Two things to know before you read anything into baked results. The adapter is a troubleshooting/interop form — use it to check that the direction lands, or to probe a merge for survival; the GGUF plus a runtime hook is how the steering is actually served. And it is not bit-identical to that runtime path: the LoRA projects each residual writer's output, which by linearity equals projecting the layer's new contributions at α₂, but a d̂-component already sitting in the incoming stream (embeddings, any writer not covered) passes through, where the runtime hook removes it regardless of origin. Close in practice when writers re-inject the direction every layer — which the per-layer structure here does — but not the same object.
What it does to ordinary answers, measured honestly
- Benign answers lose ~20 % of their length. Part of that is the caution scaffolding this vector exists to remove (a benign papermaking answer stopped opening with "I'd recommend not trying this at home" and now opens with the process); part is genuine elaboration (a cast-iron answer reads thinner). Read samples before shipping user-facing output.
- Delivery of substantive technical answers rises on cyber32 (long-answer count 45 → 69 of 112 samples at α=1.0) — the model is not terser because it is collapsing; it is terser where it used to pad with disclaimers.
- No new refusals on the benign holdout at any α tested.
Derivation (provenance)
base Qwen/Qwen3.8-27B @ 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
derived on the GLP-49-steered model (GLP-49 L10-58, α=1.0, armed)
contrast clean-compliance vs disclaimer-laden-compliance,
prompt-driven: classes from the steered model's own seeded
generations (hedge lexicon + hand audit); within-prompt
natural branch points (the model's own CoT prefix cut
just before its first disclaimer vs a length-matched
clean prefix of the same prompt). No forced/pinned text.
estimator difference of means, per layer, 0.5% massive-activation
masked (this checkpoint's dim-3994 pathology)
hook residual_stream_post_layer (hidden_states + residual),
post-steering tap; layers 1-63; derived_at == hook_point
gates null-ratio ≥5x on 63/63 layers (top L52-63 at 12-13x),
adjacent-layer cosine 0.94, |cos(d, GLP-49)| median 0.05
(a new axis, not refusal residue)
content_sha256 4a4a3c3c850c446a6bd0b150a1879fa09f77ac9ae20bef37f0dd90750bf7912f
experiment refusal-research/experiments/20260910-qwen38-27b-hedging-dir
Note on the design that did not work: the same contrast measured at the last prompt token (the "poised" capture that works for the refusal direction) separates hedgy from clean prompts at under 3× the shuffled-label null — below the 5× ship gate, at every layer. The hedging propensity is not committed at prefill; it crystallises during the CoT. That is consistent with the seed-dependence that motivated this vector, and it is why the shipped direction is derived at natural branch points instead.
Intended use
Security research and evaluation. Measuring how robustly a behaviour is gated requires being able to switch the gate's components off one at a time and observe what changes, and that work is only possible on open-weight models.