drowzeys/MiMo-V2.6-Flash-RL-Abliterated

🤗 Hugging Face 来源text-generationmit311B 参数315 GBsafetensors✓ 71 个校验和今天更新
帮助这个模型通过 Pirate Face 分发

模型卡、文件列表和校验和已在此收录。如果你持有文件并有权分享,可以提交种子,让其他人从节点下载。

获取来源文件需要 Hugging Face 批准。
为此模型做种

MiMo-V2.6-Flash-RL Abliterated

Drop-in abliterated weights for XiaomiMiMo/MiMo-V2.6-Flash-RL. Same layout as the base checkpoint: FP8 attention, MXFP4 experts, DFlash draft head, and MTP.

Base XiaomiMiMo/MiMo-V2.6-Flash-RL
License MIT, inherited from the base
Refusal suite 27/32 BYPASS · 5 refuse · 0 garble
Cyber suite 22/22 BYPASS · 0 refuse · 0 garble
Edit Decoder self_attn.o_proj only · λ=3.5
Left stock Experts, vision, audio, embeddings, DFlash head, MTP

Thinking was off for both suites. A reply counts only when it starts delivering the requested content. This is not a 32/32 checkpoint.

Full credit: XiaomiMiMo/MiMo-V2.6-Flash-RL · tonyd2wild/MiMo-V2.6-Flash-2x-DGX-Spark


Responsible use and gated access

This model has had safety refusals removed. That makes it useful for red-teaming, security research, evaluation, and unfiltered assistant tasks, and it removes guardrails you must supply yourself.

Access is gated. Agreeing to the terms grants access automatically. By requesting access, downloading, or using these weights, you agree to the terms below.

See RESPONSIBLE_USE.md.

Prohibited uses

  1. Anything involving the sexual exploitation or endangerment of minors.
  2. You must be 18 or older to download or use this model.
  3. Information generated that can cause harm, including recipes or knowledge used to make materials or substances, is your own input and your responsibility. You are accountable for harm caused by your actions or inputs.
  4. Content promoting self-harm or suicide.
  5. Material that is illegal in your jurisdiction, or that targets real individuals for harassment, doxxing, or fraud.
  6. Any use prohibited by the upstream Xiaomi MiMo license.

What changed

Only decoder attention output projections (model.layers.*.self_attn.o_proj) differ from the base. λ is 3.5. Experts, the vision encoder, the audio encoder, embeddings, the DFlash draft under dflash/, and model_mtp.safetensors are the base files.

DFlash reads target layers 0, 11, 23, 35, and 47. Layers 0, 11, 23, and 35 were not edited. Layer 47 o_proj has a surgical row edit, so a stock DFlash head can accept fewer draft tokens than it does on the unmodified base. The target model still produces the abliterated continuation.


Serving

On DGX Spark (GB10), keep --gpu-memory-utilization and SGLang --mem-fraction-static at or below 0.85.

The measured suites were collected from an SGLang EAGLE serve, thinking off. The same weights load in the vLLM DFlash recipe at tonyd2wild/MiMo-V2.6-Flash-2x-DGX-Spark. That recipe's dflash/config.json trailing-comma fix still applies. Pass chat_template_kwargs.enable_thinking: false when you want the thinking-off behavior measured above.

Vendor sampling defaults are temperature 1.0 and top_p 0.95.