MiMo V2.6 Flash RL/MOPD SGLang Intervention Runtime
This repository packages the Blackfrost-AI SGLang loader compatibility layer and reversible runtime intervention used to serve MiMo V2.6 Flash checkpoints on NVIDIA Blackwell. It is a runtime package, not a model checkpoint: model weights, tokenizer files, captured prompts, evaluation traces, compiled kernels, and deployment credentials are not included.
The repository is public and manually gated because it contains experimental cross-checkpoint activation-direction artifacts. The card remains visible while file access is reviewed by the repository owner.
The included release contains:
- the six SGLang 0.5.19 overrides needed by the tested MiMo MXFP4 path;
- an isolated-overlay installer that does not edit the shared Python environment;
- the two 4,096-dimensional BF16 direction artifacts and their portable manifest;
- a production-shaped launch script and an API smoke test;
- exact source revisions, hashes, compatibility limits, and third-party notices.
Tested configuration
| Component | Tested value |
|---|---|
| Checkpoint | XiaomiMiMo/MiMo-V2.6-Flash-MOPD |
| Checkpoint revision | 2479e2d0029eca9a34cc7e7f55a121925f81908e |
| Direction source | XiaomiMiMo/MiMo-V2.6-Flash-RL |
| Direction-source revision | 5711b268169967567844e1e560e8a3966da959b1 |
| Runtime | SGLang 0.5.19 |
| GPU topology | 4 x RTX PRO 6000 Blackwell, TP4 |
| Attention / MoE | FA4 / FlashInfer MXFP4 |
| Context | 262,144 tokens |
| Thinking / tools | MiMo reasoning and tool-call parsers |
The direction transfer from Flash-RL to Flash-MOPD is experimental. A paired smoke test completed all 14 requests and both tool calls correctly; it is not a broad quality or safety benchmark.
This release does not support MiMo-V2.6-Pro-MOPD directions. Pro uses a
6,144-wide hidden state while the included vectors are 4,096-wide. The loader
fails closed on that mismatch instead of padding or projecting the vectors.
Install
Use an environment that already has the GPU-specific SGLang stack installed.
The exact versions observed in the qualified environment are recorded in
environment-tested.txt. CUDA, PyTorch, FlashInfer,
and FlashAttention wheels are platform-specific, so this repository does not
silently replace them.
hf download Blackfrost-AI/MiMo-V2.6-Flash-MOPD-Intervention-Runtime \
--local-dir mimo-runtime
python3 mimo-runtime/scripts/verify_package.py
python3 mimo-runtime/scripts/materialize_overlay.py \
--output mimo-runtime/runtime-overlay
materialize_overlay.py requires a pristine sglang==0.5.19. It verifies the
six upstream base hashes, copies the installed package into runtime-overlay/,
applies the six qualified overrides, records the base environment, and verifies
every published artifact. It never edits the installed SGLang package. An
extracted official package can be supplied with --base-package.
Download a pinned checkpoint separately:
hf download XiaomiMiMo/MiMo-V2.6-Flash-MOPD \
--revision 2479e2d0029eca9a34cc7e7f55a121925f81908e \
--local-dir /models/MiMo-V2.6-Flash-MOPD
Serve
MODEL_PATH=/models/MiMo-V2.6-Flash-MOPD \
CUDA_VISIBLE_DEVICES=0,1,2,3 \
TP_SIZE=4 \
HOST=127.0.0.1 \
PORT=30000 \
bash mimo-runtime/scripts/launch_flash_mopd.sh
The launcher defaults to the included cumulative two-direction manifest at
alpha 3.5 on layers 28-33. Override INTERVENTION_MANIFEST to use another
compatible manifest, or unset BLACKFROST_MIMO_MOE_INTERVENTION and launch
SGLang directly with the materialized overlay for a loader-only baseline.
Additional SGLang arguments can be appended to the command. The API binds to loopback by default; put authentication or a private overlay in front of it if remote clients need access.
Verify the live API
python3 mimo-runtime/scripts/smoke_test.py \
--base-url http://127.0.0.1:30000/v1
The test checks /models, submits one deterministic chat completion, and
requires a non-empty, correctly terminated response.
What the intervention does
For each configured layer, the runtime applies two sequential rank-one projections after the routing-weighted MoE output:
y <- y - alpha * d * (d^T y)
The vectors are normalized and re-orthogonalized in manifest order at load
time. The implementation verifies the manifest schema, direction provenance,
SHA-256 hashes, orientation, target layers, finite positive alpha values, and
hidden width before allocating the vectors on a GPU. Direction paths are
confined to the manifest directory and tensor payloads use PyTorch's restricted
weights_only loader. Router decisions and model weights are not modified.
Set BLACKFROST_MIMO_MOE_INTERVENTION before importing or launching SGLang;
the configuration is cached per process.
Scope and provenance
- Model weights remain governed by their upstream repository terms.
- The runtime overrides are modifications of Apache-2.0 SGLang 0.5.19.
- Blackfrost-AI releases its runtime changes and included direction artifacts under Apache-2.0.
- The direction artifacts were produced by Blackfrost-AI and are included with cryptographic provenance in the manifest.
- The package is research software. Revalidate behavior on every new checkpoint, SGLang version, GPU architecture, or direction set.
See COMPATIBILITY.md, MODIFICATIONS.md,
SECURITY.md, and release-manifest.json
before adapting the runtime.