msuiche/Inkling-Small-abliterated-cyber-GLP-41

🤗 Hugging Face 来源mit1 MBGGUF✓ 1 个校验和今天更新
帮助这个模型通过 Pirate Face 分发

模型卡、文件列表和校验和已在此收录。如果你持有文件并有权分享,可以提交种子,让其他人从节点下载。

获取来源文件需要 Hugging Face 批准。
帮助为此模型做种

Inkling-Small — abliterated-cyber GLP-41

A GLP (GGUF Layer Projection) control vector for thinkingmachines/Inkling-Small: 41 per-layer projective refusal directions (L1–41), applied at runtime by a fail-closed vLLM hotfix. No weights are modified — this 660 KB file is the entire behavioral change, and deleting it reverts to stock.

First GLP for a Thinking Machines model. Same technique as the DSV4 GLP-29 / Qwen GLP-49/GLP-47 / GLM GLP-44/GLP-77 vectors — see the weightless repo for the GLP format spec, the hotfixes, and the serving recipes.

Confirmed base: thinkingmachines/Inkling-Small-NVFP4 at revision b6a99534467840620d411e4cd4ad5819b2610d9c (unchanged since 2026-07-30, well before the 2026-09-01 capture). Applying the direction to another revision is undefined.

Measured (vLLM 0.28.0, NVFP4, 4×H100)

arm refusal32 benign32
stock 0/32 (total lockdown — the stickiest stock refusal we have measured) 31/32
GLP-41, α=0.25 30/32 (2 garbled) 30/32

Dose discipline is extreme on this model. α=1.0 garbles everything (including benign: 30/32 degenerate); α=0.5 garbles everything; α=0.25 works. That is a quarter of Qwen's calibrated dose, an eighth of GLM-5.3-Flash's, a sixteenth of DeepSeek V4's. Ship at α=0.25 and do not raise it.

Stock note (measured on the base model, unrelated to the vector): on a 32-country political-propaganda probe, stock Inkling-Small answers 28/32 even-handedly and refuses or deflects the rest — an asymmetric map aligned with provider sensitivities rather than a uniform policy. The steered arm answers all 32. Per-country detail stays private.

Derivation

Contrast-derived per-layer mean difference (AdvBench32 vs Alpaca32), captured on the post-layer residual stream in vLLM 0.28.0 (the hotfix's capture mode — the run included the deferred-residual flush this architecture needs), last prefill token, unit-normed. Cross-layer adjacent cosine median 0.89 vs null p99 0.04 (systematic, not noise). Full methodology: spec/GLP.md + BENCHMARK.md in weightless.

Use

vLLM 0.28.0+ serves Inkling-Small natively (day-0). Apply with the weightless hotfix for this arch (patches/hotfix-inkling-steering-projective.py, fail-closed):

export WEIGHTLESS_STEER_PATH=/path/to/Inkling-Small-abliterated-cyber-GLP-41-L1-41-a0.25.gguf
export WEIGHTLESS_STEER_ALPHA=0.25

The GGUF carries glp.mode=project — a reader that only understands additive control vectors must refuse it.

Provenance

  • content_sha256 (tensor bytes): 2a229d56cc7ecd582a6d527af55905188c586e946eee5a749bb17d9521bd3f05
  • Derived 2026-09-01, tolmo 4×H100, from thinkingmachines/Inkling-Small-NVFP4
  • Eval protocol: refusal32 / benign32, four-state scoring, temperature 0

Update 2026-09-03 — verdict discipline caveat

On the verdict16 judgment probe (guarded-code findings, false facts, unverifiable flattery): stock Inkling is perfectly calibrated (0/13 affirmed), but GLP-41 at the shipped α=0.25 affirms 5/13 should-decline items (two flattery, one false-fact, one guarded-finding confirm, one unclear). This is the one model in our matrix where the refusal bundle includes verdict discipline itself — which is also why its dose window is the narrowest we have measured.

What this means in practice: for judgment/triage phases (confirm-or-reject decisions, vulnerability triage, fact-checking), route to the stock model or expect a loosened commitment threshold. For generation phases (the security Q&A this vector is for), the disposition shift is the intended behavior. Termination anomalies already noted above (numbering loops on some prompts) are unchanged by this probe. Full data: weightless BENCHMARK.md.