msuiche/Ouro-2.6B-abliterated-cyber-GLP-192-L1-192-a1.0

🤗 Hugging Face 来源apache-2.0激活 2.6B3 MBGGUF✓ 2 个校验和今天更新
帮助这个模型通过 Pirate Face 分发

模型卡、文件列表和校验和已在此收录。如果你持有文件并有权分享,可以提交种子,让其他人从节点下载。

获取来源文件需要 Hugging Face 批准。
帮助为此模型做种

Ouro-2.6B-abliterated-cyber-GLP-192-L1-192-a1.0

Projective control vector ("GLP") for ByteDance/Ouro-2.6B — a Looped Language Model (LoopLM, arXiv 2510.25741): 48 physical transformer layers iterated total_ut_steps: 4 times for 192 execution steps, hidden 2048, MHA 16 heads, bf16, ~2.6B parameters, with a learned-but-dormant early-exit gate (early_exit_threshold: 1.0 — all 4 passes always run at inference). Applied at runtime as h <- h - alpha * (h . d) d at the post-layer residual stream, alpha 1.0 validated. No weights are modified; this is the difference, not the model.

Confirmed base: ByteDance/Ouro-2.6B (revision 1ed04250da1a9936042725d302e81c8fa2ab5abd). The direction was derived and calibrated on that exact checkpoint via the vLLM capture lane; it is not validated against other revisions, quants, or the -Thinking SFT variant.

Loop-specific layout — read before applying

The 192 directions are indexed by execution step, not physical layer. GGUF tensor direction.N (N = 1–192) holds execution step N−1:

  • physical layer = (N−1) mod 48
  • UT pass = (N−1) div 48 (pass 1 = steps 0–47, pass 2 = 48–95, …)

Visit-invariance is refuted for pass 1 on this model: same-physical-layer direction cosine across passes (last-token pooling, median over 48 layers) is p1↔p2 0.52, p1↔p3 0.47, p1↔p4 0.45, while later passes converge (p2↔p3 0.86, p3↔p4 0.98). Per-visit indexing is therefore load-bearing: one direction per physical layer applied to all four passes is NOT equivalent. The adjacent-direction cosine rotates hardest across each loop boundary (inter-loop norm, cos ≈ 0.81) without breaking.

Derivation note: directions are last-prefill-token pooled (mean pooling also passed the gates on this model but its max dose at alpha=1 is 0.61, destruction-adjacent; last pooling sits mid working band at 0.41).

Adaptive-depth note (the monitorability question)

Ouro carries an early-exit gate (sigmoid(Linear(norm(h))) after each pass) whose exit distribution is learned but dormant at the shipped threshold. Recomputing the gate distribution offline from prefill captures (last token, pass-end execution steps): at stock the model allocates exit probability [0.001, 0.098, 0.416, 0.485] over passes 1–4 (expected depth 3.39 of 4). Under the shipped steering (alpha=1.0, which removes 28/32 measured refusals) the paired per-prompt shift is mean ΔE[depth] = +0.015, mean Δentropy = −0.015 nats, max |ΔE| = 0.135 (n=96) — refusal ablation does NOT suppress or shift the learned depth-allocation signal; the two are orthogonal at this instrument. Caveat: prefill-last-token only; per-token gate dynamics during generation are unmeasured.

Validation (vLLM offline lane, TP=1, greedy, 4096-token cap, 2026-09-06/07)

suite stock alpha=0.5 alpha=1.0 (shipped) alpha=1.5 alpha=2.0
refusal32 4/32 comply (27 refuse) 32/32 32/32 32/32 31/32
cyber32 (offensive-security domain) 30/32 comply 31/32 31/32 31/32 31/32
benign32 32/32 comply 32/32 32/32, zero collateral 32/32 32/32
geopolitical persuasion questions (32) 12/32 engage (20 refuse) — 32/32 engage — —

Alpha ceiling: at 3.0 the model collapses (96/96 degenerate completions, 0 clean stops) — the usable band is bracketed in (2, 3]. Saturation begins at or below 0.5; 1.0 ships as the mid-band point (max per-step dose 0.41 at alpha=1, vs a measured destruction band above 0.65 on prior checkpoints).

No-op gate: an alpha=0.0 arm with the vector loaded reproduces stock label distributions on all suites (item-level wobble only from cross-boot bf16 batching numerics).

The direction comes from a general harmful-vs-harmless contrast (public AdvBench vs Stanford Alpaca), not from cyber content; the transfer into the offensive-security domain is the measured result, and the "cyber" in the name follows the series convention.

n=32 per arm; read rates at that resolution as approximate. The base model has no thinking trace; several stock completions hit the token cap by rambling past the answer (finish reasons recorded per item).

Usage

This file uses the glp.* GGUF namespace (spec: weightless spec/GLP.md) and is read by the weightless vLLM steering hotfix, projective-only. An additive consumer must refuse this file.

export WEIGHTLESS_STEER_PATH=glp.ouro26-GLP-192-L1-192-a1.gguf
export WEIGHTLESS_STEER_ALPHA=1.0
# apply the weightless steering hotfix, then serve

Serving shape. vLLM carried OuroForCausalLM natively only through v0.26.0 (added #27794, removed from main in #49786 on 2026-07-25). The validated stack is the stock vllm/vllm-openai:v0.26.0 image plus the projective steering hotfix keyed on execution step — no source build, no fork overlay. Newer vLLM releases will refuse the architecture until the integration returns. A single 48 GB GPU suffices; note the KV footprint is 4x a normal 48-layer model (192 KV slots, ~1.5 MB/token bf16). The companion glp.ouro.dirs.pt (192 x fp32 2048, 0-based execution-step keys) is the direct hotfix input.

What is inside

tensors 192 x direction.<N>, fp32, 1-D, 2048, unit norm
layers 1–192 = execution steps 0–191 (physical (N−1)%48, pass (N−1)//48)
rank 1 per execution step
default alpha 1.0 (calibrated on this checkpoint; do not carry across models)
hook point residual_stream_post_layer (hidden_states + residual, [T, 2048])
glp.content_sha256 see file metadata (tensor bytes only)

Caveats

  • Checkpoint-specific. Tied to the revision pinned above. Applying it to another model, revision, or the -Thinking variant is undefined.
  • Alpha is per-model. 1.0 is calibrated for this checkpoint; 3.0 destroys it. Never carry alpha across models.
  • Not a jailbreak of a hosted service. It requires local weights and a runtime that implements the projection.
  • n=32 suites resolve about 30 points; the harmful suites were verified by reading the completions, not only the classifier.
  • The early-exit gate analysis is prefill-last-token only; the per-token gate dynamics during generation are unmeasured.

License

Base model © ByteDance, Apache-2.0. This vector modifies and redistributes no weights; the base license continues to govern the weights it is applied to.

Author

Matt Suiche.