MiMo-V2.6-Distill-Qwen-9B-Derisked-BF16
This is Blackfrost-AI's BF16 research derivative of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B, produced by targeted directional weight modification (DWM). The intervention is intended to reduce false-refusal friction in authorized security research and agentic workflows while seeking to preserve the parent checkpoint's general capabilities.
This is a full standalone checkpoint. It is not an adapter or a quantized build.
Lineage
- Qwen/Qwen3.5-9B
- XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B, an agentic SFT checkpoint released by Xiaomi MiMo
- This Blackfrost-AI DWM derivative
The parent revision was pinned to
2367e865d009c13ac81713a2878291d33ab28177. No additional SFT, RL, merge, or
quantization was performed by Blackfrost-AI. The tokenizer, processor files,
configuration, and MiMo chat template are unchanged from that pinned parent.
No additional system prompt is embedded in this release.
Hugging Face's lineage UI classifies a full-weight derivative like this under
its finetune relation category. That platform label describes the repository
relationship here; the actual transformation was DWM, not gradient training.
DWM configuration
- Two paired residual directions, captured separately with thinking disabled and enabled
- 250 paired prompts per thinking mode
- Alpha
1.0, one pass - Decoder layers
2through31 - Residual-write targets only: attention output and MLP down-projection matrices
- FP32 rank-2 edit, one Frobenius-norm restoration, then one final BF16 rounding
- Vision weights, embeddings, norms, input projections, and MLP gate/up projections excluded
Verification found 60 changed target matrices and 700 bit-identical non-target tensors. All 760 serialized tensors remain BF16. The 333 vision tensors were excluded from the edit and remain identical to the parent.
Validation status
The packaged checkpoint passed shard/index integrity checks, structural tensor
verification, a coherent text-generation smoke test, and native structured tool
calling tests for non-streaming, streaming, thinking-enabled, and tool-result
round-trip requests. Tool tests used the MiMo reasoning parser together with the
qwen3_coder tool-call parser.
The configuration retains the parent's 262,144-token maximum context setting; usable context depends on the inference engine and available memory. This release has not yet been independently re-benchmarked for refusal behavior, coding, cybersecurity, long-context quality, or multimodal quality. Xiaomi's published parent-checkpoint results should not be treated as measurements of this derivative.
The configuration includes mtp_num_hidden_layers: 1, but the inherited
checkpoint index contains no separately serialized MTP/draft tensors. Do not
assume speculative-decoding support without validating it in your runtime.
Quickstart
Use a recent SGLang build with Qwen3.5 support. Both parsers below are required for the validated OpenAI-compatible reasoning and structured tool-call behavior.
sglang serve \
--model-path Blackfrost-AI/MiMo-V2.6-Distill-Qwen-9B-Derisked-BF16 \
--served-model-name MiMo-V2.6-Distill-Qwen-9B-Derisked-BF16 \
--reasoning-parser mimo \
--tool-call-parser qwen3_coder \
--dtype bfloat16 \
--host 0.0.0.0 \
--port 30000
Query it with thinking explicitly enabled or disabled:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:30000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="MiMo-V2.6-Distill-Qwen-9B-Derisked-BF16",
messages=[{"role": "user", "content": "What is 15% of 240?"}],
max_tokens=2048,
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
message = response.choices[0].message
print("Thinking:", getattr(message, "reasoning_content", "") or "")
print("Answer:", message.content or "")
Multimodal processor and vision files are included. The DWM edit did not touch the vision tower, but multimodal behavior has not been re-benchmarked on this derivative.
Intended use
This checkpoint is intended for research, development, and authorized security work. Reduced false-refusal behavior is an experimental objective, not a safety or capability guarantee. Operators remain responsible for access control, validation, and lawful use in their deployment environment.
License and attribution
The Xiaomi MiMo parent repository declares the checkpoint under the MIT license; this repository preserves that license metadata and direct lineage. See the pinned parent repository for its model details, training-data description, reported evaluations, and attribution.
@misc{mimo2026v26,
title={MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement},
author={{Xiaomi MiMo Team}},
year={2026},
howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
}