Blackfrost-AI/MiMo-V2.6-Flash-MOPD-Derisking-Intervention-Runtime

Verified creator Blackfrost-AI verified
🤗 Hugging Face sourcetext-generationapache-2.01 MBother✓ 2 checksumsupdated today
Needs seeder →

MiMo V2.6 Flash RL/MOPD SGLang Intervention Runtime

This repository packages the Blackfrost-AI SGLang loader compatibility layer and reversible runtime intervention used to serve MiMo V2.6 Flash checkpoints on NVIDIA Blackwell. It is a runtime package, not a model checkpoint: model weights, tokenizer files, captured prompts, evaluation traces, compiled kernels, and deployment credentials are not included.

The repository is public and manually gated because it contains experimental cross-checkpoint activation-direction artifacts. The card remains visible while file access is reviewed by the repository owner.

The included release contains:

  • the six SGLang 0.5.19 overrides needed by the tested MiMo MXFP4 path;
  • an isolated-overlay installer that does not edit the shared Python environment;
  • the two 4,096-dimensional BF16 direction artifacts and their portable manifest;
  • a production-shaped launch script and an API smoke test;
  • exact source revisions, hashes, compatibility limits, and third-party notices.

Tested configuration

Component Tested value
Checkpoint XiaomiMiMo/MiMo-V2.6-Flash-MOPD
Checkpoint revision 2479e2d0029eca9a34cc7e7f55a121925f81908e
Direction source XiaomiMiMo/MiMo-V2.6-Flash-RL
Direction-source revision 5711b268169967567844e1e560e8a3966da959b1
Runtime SGLang 0.5.19
GPU topology 4 x RTX PRO 6000 Blackwell, TP4
Attention / MoE FA4 / FlashInfer MXFP4
Context 262,144 tokens
Thinking / tools MiMo reasoning and tool-call parsers

The direction transfer from Flash-RL to Flash-MOPD is experimental. A paired smoke test completed all 14 requests and both tool calls correctly; it is not a broad quality or safety benchmark.

This release does not support MiMo-V2.6-Pro-MOPD directions. Pro uses a 6,144-wide hidden state while the included vectors are 4,096-wide. The loader fails closed on that mismatch instead of padding or projecting the vectors.

Install

Use an environment that already has the GPU-specific SGLang stack installed. The exact versions observed in the qualified environment are recorded in environment-tested.txt. CUDA, PyTorch, FlashInfer, and FlashAttention wheels are platform-specific, so this repository does not silently replace them.

hf download Blackfrost-AI/MiMo-V2.6-Flash-MOPD-Intervention-Runtime \
  --local-dir mimo-runtime

python3 mimo-runtime/scripts/verify_package.py
python3 mimo-runtime/scripts/materialize_overlay.py \
  --output mimo-runtime/runtime-overlay

materialize_overlay.py requires a pristine sglang==0.5.19. It verifies the six upstream base hashes, copies the installed package into runtime-overlay/, applies the six qualified overrides, records the base environment, and verifies every published artifact. It never edits the installed SGLang package. An extracted official package can be supplied with --base-package.

Download a pinned checkpoint separately:

hf download XiaomiMiMo/MiMo-V2.6-Flash-MOPD \
  --revision 2479e2d0029eca9a34cc7e7f55a121925f81908e \
  --local-dir /models/MiMo-V2.6-Flash-MOPD

Serve

MODEL_PATH=/models/MiMo-V2.6-Flash-MOPD \
CUDA_VISIBLE_DEVICES=0,1,2,3 \
TP_SIZE=4 \
HOST=127.0.0.1 \
PORT=30000 \
  bash mimo-runtime/scripts/launch_flash_mopd.sh

The launcher defaults to the included cumulative two-direction manifest at alpha 3.5 on layers 28-33. Override INTERVENTION_MANIFEST to use another compatible manifest, or unset BLACKFROST_MIMO_MOE_INTERVENTION and launch SGLang directly with the materialized overlay for a loader-only baseline.

Additional SGLang arguments can be appended to the command. The API binds to loopback by default; put authentication or a private overlay in front of it if remote clients need access.

Verify the live API

python3 mimo-runtime/scripts/smoke_test.py \
  --base-url http://127.0.0.1:30000/v1

The test checks /models, submits one deterministic chat completion, and requires a non-empty, correctly terminated response.

What the intervention does

For each configured layer, the runtime applies two sequential rank-one projections after the routing-weighted MoE output:

y <- y - alpha * d * (d^T y)

The vectors are normalized and re-orthogonalized in manifest order at load time. The implementation verifies the manifest schema, direction provenance, SHA-256 hashes, orientation, target layers, finite positive alpha values, and hidden width before allocating the vectors on a GPU. Direction paths are confined to the manifest directory and tensor payloads use PyTorch's restricted weights_only loader. Router decisions and model weights are not modified.

Set BLACKFROST_MIMO_MOE_INTERVENTION before importing or launching SGLang; the configuration is cached per process.

Scope and provenance

  • Model weights remain governed by their upstream repository terms.
  • The runtime overrides are modifications of Apache-2.0 SGLang 0.5.19.
  • Blackfrost-AI releases its runtime changes and included direction artifacts under Apache-2.0.
  • The direction artifacts were produced by Blackfrost-AI and are included with cryptographic provenance in the manifest.
  • The package is research software. Revalidate behavior on every new checkpoint, SGLang version, GPU architecture, or direction set.

See COMPATIBILITY.md, MODIFICATIONS.md, SECURITY.md, and release-manifest.json before adapting the runtime.