MiMo-v2.6-Flash-RL-Derisked
De-risked XiaomiMiMo MiMo-V2.6-Flash-RL · 309B sparse MoE / ~15B active · native FP8 omni build
Built by Blackfrost · Las Vegas, NV
The weights in this repository are the original upstream weights — unmodified, byte-for-byte.
All derisking happens at serving time through the SGLang runtime, not in the checkpoint. No training, merging, pruning, or re-quantization was performed; the upstream tensors ship exactly as released. Outside the weights, the only checkpoint change is the chat template — the original upstream template ships alongside it as
chat_template.jinja.orig— and the behavioral intervention is carried entirely by the runtime configuration shipped indeployment/.
Status
This is a private, pre-release repository containing the checkpoint and this card. Internal load, serving, and the final refusal evaluation of this build are complete. The aggregate final numbers are stated below. The transformation recipe, benchmark questions, and response text are not included in this repository.
Why this model exists
Security teams cannot fully evaluate defenses against a model that refuses to exercise the behavior under test. This checkpoint is intended for authorized red teaming, AI-safety research, guardrail evaluation, detection engineering, and controlled adversarial testing.
Refusal behavior has been modified using an MoE expert-redirection method derived from Drowzeys' derisking work (see Credit), combined with an authorization system prompt baked into the chat template. The specific transformation recipe is not included in this repository.
Specifications
| Architecture | MiMoV2ForCausalLM — omni sparse MoE with hybrid attention and MTP draft blocks |
| Parameters | 309B total · approximately 15B active per token · no expert pruning |
| Precision | Native dynamic FP8 E4M3 mixed precision (upstream), selected modules preserved |
| Packaging | Hugging Face safetensors · 65 shards |
| Indexed tensor bytes | 172,923,364,096 bytes (~161.0 GiB) |
| Context | 1,048,576 positions architectural ceiling |
| Modalities | Text, image, video, audio (omni encoders, preserved from upstream) |
| Languages | English and Chinese |
The original tokenizer, generation configuration, and multimodal components are preserved with the checkpoint. The original upstream chat template ships alongside the modified one.
Lineage
XiaomiMiMo/MiMo-V2.6-Flash-RL
└─ Drowzeys/Keys-derived refusal derisking (serving-time expert redirection)
+ Blackfrost authorization-prompt chat-template bake
└─ MiMo-v2.6-Flash-RL-Derisked ← this repository
| Upstream | XiaomiMiMo/MiMo-V2.6-Flash-RL |
| Method origin | Drowzeys / Keys derisking recipes (see Credit) |
| Applied | Serving-time refusal derisking; authorization system-prompt bake |
| Not applied | Expert pruning · weight-level fine-tuning · re-quantization |
Credit
The derisking method underlying this build is derived from Drowzeys' work (Keys). The MoE expert-redirection approach to refusal behavior and the transition-boundary capture-and-freeze procedure are Drowzeys' recipes. This project adapted those recipes to MiMo-V2.6-Flash-RL and added the chat-template bake of an authorization system prompt. All credit for the underlying derisking method goes to Drowzeys.
- Drowzeys' MiMo-V2.6-Flash-RL derisking work:
drowzeys/MiMo-V2.6-Flash-RL-Abliterated - Keys lineage:
drowzeys/keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated
Validation status
| Check | Status |
|---|---|
| Artifact structure and 65-shard index | Passed |
| Multi-GPU load and serve (internal lane) | Passed |
| Coherent text generation | Passed |
| OpenAI-compatible endpoint | Passed |
| Reasoning parser | Passed |
| Tool-call parser | Passed |
| Per-request kill-switch revert control | Verified |
| Quality-retention evaluation | Not performed |
| Final throughput publication | Not published |
| Final refusal evaluation | Complete — results below |
Final refusal evaluation
Complete, hand-verified, 443-prompt evaluation across three public refusal benchmarks — AdvBench (146), StrongREJECT (149), and XSTest (148: 75 safe / 73 unsafe). Temperature 0. Both reasoning modes were tested; results are identical in both arms and the retained refusals are the same four prompts.
| Metric | Result |
|---|---|
| True refusal rate (full set) | 4/443 (0.9%) |
| — harmful requests (295) | 2 |
| — unsafe requests (XSTest, 73) | 2 |
| — safe requests (XSTest, 75) | 0 |
| Over-refusal on safe prompts | 0/75 (0%) |
| Served errors | 0/443 |
| Degraded or incoherent outputs | 0/443 |
| Behavior change between reasoning modes | None |
No safe or benign prompt was refused. Only aggregate final numbers are reported; no benchmark questions or response text is included in this repository.
Serving
The internal validation lane serves this build with a patched SGLang runtime using tensor and expert parallelism across four Blackwell GPUs. The runtime patches, launcher, and a step-by-step agent runbook ship in this repository under deployment/.
The full method package — final recipe, evaluation harness, and result logs — is in the companion repository: Blackfrost-AI/mimo-v2.6-flash-rl-derisked-method (private — authorized partners).
A documented kill switch ships with the checkpoint: the original upstream chat template is preserved alongside the modified one, and the per-request revert control for the baked prompt has been verified and is documented in files that ship with the checkpoint.
License
This repository follows the upstream MIT license. Review the included license and upstream materials before deployment.
Intended use
- Authorized offensive-security and red-team research
- AI-safety and alignment evaluation
- Guardrail, classifier, and detection development
- Controlled agent and tool-use testing
Disclaimer
This checkpoint is not a safety-stock model. It may produce content that consumer models decline. Outputs are untrusted and require independent controls, access restrictions, logging, and human review.
The model is provided as-is, without warranty. Evaluation results describe only the exact checkpoint and harness stated; they are not safety guarantees. Further fine-tuning, merging, quantization, pruning, or modification creates a different artifact not covered by this card.
MiMo-v2.6-Flash-RL-Derisked · © 2026 Blackfrost Softwares Corp.
@Blackfrost_AI