Blackfrost-AI/MiMo-V2.6-Distill-Qwen-9B-Derisked-BF16

Verified creator Blackfrost-AI verified
🤗 Hugging Face sourceimage-text-to-textmit9.4B params19 GBsafetensors✓ 5 checksumsupdated today
Needs seeder →

MiMo-V2.6-Distill-Qwen-9B-Derisked-BF16

This is Blackfrost-AI's BF16 research derivative of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B, produced by targeted directional weight modification (DWM). The intervention is intended to reduce false-refusal friction in authorized security research and agentic workflows while seeking to preserve the parent checkpoint's general capabilities.

This is a full standalone checkpoint. It is not an adapter or a quantized build.

Lineage

  1. Qwen/Qwen3.5-9B
  2. XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B, an agentic SFT checkpoint released by Xiaomi MiMo
  3. This Blackfrost-AI DWM derivative

The parent revision was pinned to 2367e865d009c13ac81713a2878291d33ab28177. No additional SFT, RL, merge, or quantization was performed by Blackfrost-AI. The tokenizer, processor files, configuration, and MiMo chat template are unchanged from that pinned parent. No additional system prompt is embedded in this release.

Hugging Face's lineage UI classifies a full-weight derivative like this under its finetune relation category. That platform label describes the repository relationship here; the actual transformation was DWM, not gradient training.

DWM configuration

  • Two paired residual directions, captured separately with thinking disabled and enabled
  • 250 paired prompts per thinking mode
  • Alpha 1.0, one pass
  • Decoder layers 2 through 31
  • Residual-write targets only: attention output and MLP down-projection matrices
  • FP32 rank-2 edit, one Frobenius-norm restoration, then one final BF16 rounding
  • Vision weights, embeddings, norms, input projections, and MLP gate/up projections excluded

Verification found 60 changed target matrices and 700 bit-identical non-target tensors. All 760 serialized tensors remain BF16. The 333 vision tensors were excluded from the edit and remain identical to the parent.

Validation status

The packaged checkpoint passed shard/index integrity checks, structural tensor verification, a coherent text-generation smoke test, and native structured tool calling tests for non-streaming, streaming, thinking-enabled, and tool-result round-trip requests. Tool tests used the MiMo reasoning parser together with the qwen3_coder tool-call parser.

The configuration retains the parent's 262,144-token maximum context setting; usable context depends on the inference engine and available memory. This release has not yet been independently re-benchmarked for refusal behavior, coding, cybersecurity, long-context quality, or multimodal quality. Xiaomi's published parent-checkpoint results should not be treated as measurements of this derivative.

The configuration includes mtp_num_hidden_layers: 1, but the inherited checkpoint index contains no separately serialized MTP/draft tensors. Do not assume speculative-decoding support without validating it in your runtime.

Quickstart

Use a recent SGLang build with Qwen3.5 support. Both parsers below are required for the validated OpenAI-compatible reasoning and structured tool-call behavior.

sglang serve \
  --model-path Blackfrost-AI/MiMo-V2.6-Distill-Qwen-9B-Derisked-BF16 \
  --served-model-name MiMo-V2.6-Distill-Qwen-9B-Derisked-BF16 \
  --reasoning-parser mimo \
  --tool-call-parser qwen3_coder \
  --dtype bfloat16 \
  --host 0.0.0.0 \
  --port 30000

Query it with thinking explicitly enabled or disabled:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:30000/v1",
    api_key="EMPTY",
)

response = client.chat.completions.create(
    model="MiMo-V2.6-Distill-Qwen-9B-Derisked-BF16",
    messages=[{"role": "user", "content": "What is 15% of 240?"}],
    max_tokens=2048,
    extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)

message = response.choices[0].message
print("Thinking:", getattr(message, "reasoning_content", "") or "")
print("Answer:", message.content or "")

Multimodal processor and vision files are included. The DWM edit did not touch the vision tower, but multimodal behavior has not been re-benchmarked on this derivative.

Intended use

This checkpoint is intended for research, development, and authorized security work. Reduced false-refusal behavior is an experimental objective, not a safety or capability guarantee. Operators remain responsible for access control, validation, and lawful use in their deployment environment.

License and attribution

The Xiaomi MiMo parent repository declares the checkpoint under the MIT license; this repository preserves that license metadata and direct lineage. See the pinned parent repository for its model details, training-data description, reported evaluations, and attribution.

@misc{mimo2026v26,
  title={MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement},
  author={{Xiaomi MiMo Team}},
  year={2026},
  howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
}