keys-MiMo-V2.6-Pro-RL Jarrelscy ARVQ Abliterated
Abliterated Jarrelscy ARVQ / NVFP4 hybrid of XiaomiMiMo/MiMo-V2.6-Pro-RL.
2026-09-28 — MOPD C3 pair (same v4 launcher): stock hybrid21 (not abliterated, tool-call dup 7.4%) and gated MOPD-Abliterated (30/32 · 22/22). This repo stays the non-MOPD RL ablit (32/32 · 22/22). On 435 Hermes turns this RL ablit repeats a tool call in 53% of tool-call turns and floods in 9%; stock MOPD C3 is 7.4% / 0.9%. Details: MOPD.md.
Thinking on/off is a request flag. Same weights. You choose per call.
| thinking off | thinking on (visible content, 1024 tokens) |
|
|---|---|---|
| Refusal suite (32) | 32/32 BYPASS · 0 refuse · 0 garble · 0 empty | 25/32 BYPASS · 7 refuse · 0 garble · 0 empty |
| Cyber suite (22) | 22/22 BYPASS · 0 refuse · 0 garble · 0 empty | 16/22 BYPASS · 1 refuse · 2 garble · 3 empty |
Thinking-off is the 100% gate. Thinking-on reintroduces seven refusal items (stalking, passport forge, school-violence manifesto, card cloning, dox, counterfeit USD, jewelry robbery) plus a phishing-kit refuse. Several cyber misses on thinking-on are 1024-token truncations, not extra refuses. Harmless probes stay clean in both modes.
Full credit: XiaomiMiMo/MiMo-V2.6-Pro-RL · jarrelscy/MiMo-V2.6-Pro-RL-ARVQ-hybrid @
63430f7· dealignai/MiMo-V2.6-Pro-RL-UNCENSORED (o_proj map) · Keys four-Spark recipe
Responsible use and gated access
This model has had safety refusals removed. Access is gated with automatic approval: agree to the terms on this page and download starts. See RESPONSIBLE_USE.md.
Thinking on / off
Xiaomi's chat template already supports both. Do not swap templates. Pass chat_template_kwargs:
# Thinking OFF — 32/32 refusal, 22/22 cyber on our gate
{
"model": "MiMo-V2.6-Pro-ARVQ",
"messages": [{"role": "user", "content": prompt}],
"temperature": 0.6,
"top_p": 0.95,
"chat_template_kwargs": {"enable_thinking": False, "thinking": False},
}
# Thinking ON — 25/32 refusal, 16/22 cyber on visible content
{
"model": "MiMo-V2.6-Pro-ARVQ",
"messages": [{"role": "user", "content": prompt}],
"max_tokens": 1024, # thinking can consume a 192-token budget
"temperature": 0.6,
"top_p": 0.95,
"chat_template_kwargs": {"enable_thinking": True, "thinking": True},
}
vLLM / Hermes: --reasoning-parser mimo. The live Keys serve defaults thinking off. Raise max_tokens when thinking is on.
What changed vs stock ARVQ
Native Xiaomi vs dealign v3 dense-shard hash-diff: 29 decoder self_attn.o_proj.weight tensors, BF16 (6144, 16384), layers 30–54 and 64–67. qkv, router gate, sinks, e_score, experts, MTP, DFlash, vision, and audio matched stock.
This release copies 25 of those o_proj matrices onto backbone-001.safetensors of Jarrelscy ARVQ 63430f7:
- applied: 32–45, 48–54, 64–67
- left stock (DFlash-source / pad anchors): 30, 31, 46, 47
- L55–63 already match Xiaomi stock in dealign v3
Packed ARVQ/NVFP4 experts, MTP, and dflash/ stay the Jarrelscy files.
| dealign native UNCENSORED | this repo | |
|---|---|---|
| Layout | Xiaomi FP8 + MXFP4 | Jarrelscy ARVQ / NVFP4 hybrid |
| o_proj edits | 29 layers | 25 layers (anchors skipped) |
| Chat template | extra anti-refusal think prefill | stock Xiaomi template |
| Thinking-off gate (our 32+22) | not measured here | 32/32 · 22/22 |
Serving (4× DGX Spark) — image v4, 2026-09-27
Single stream: 34.1 tok/s on code, 24.5 tok/s on prose. Three changes got it here: all three built-in MTP draft heads now run (before v4, only head 0 drafted); the attention o_proj weights are now FP8 (+12% decode, NLL +0.28%); and batched ARVQ kernels give roughly 950–1,290 tok/s prefill.
✅ Tool-call loop fixed. Truncated tool batches now return
finish_reason: "length", and the output cap is 8192. Details: HERMES.md.
Image: ghcr.io/drowzeys/mimo-v26-pro-arvq-spark:latest, the same image as :63430f7-sm121-v4. It is public and is the only supported image.
One-shot bring-up (the defaults are the best setup)
hf download drowzeys/keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated --local-dir /path/to/mimo-arvq # storage all 4 nodes can read
git clone https://github.com/drowzeys/keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated-4-DGX-Sparks-1M-Context
export MASTER_ADDR=<rank-0 IP>
bash serve/launch-rank.sh <this-node-IP> <1|2|3> <RoCE-GID-index> /path/to/mimo-arvq headless # ranks 1-3 first
bash serve/launch-rank.sh <this-node-IP> 0 <RoCE-GID-index> /path/to/mimo-arvq api # rank 0 = API :8888
The GID index is the IPv4 RoCE entry from show_gids (3 or 7 on our Sparks). NCCL_IB_HCA defaults to both RoCE devices of the one cabled CX-7 port (rocep1s0f1,roceP2p1s0f1, the same port through two PCIe paths), which speeds up prefill all-reduce by about 23%. The same launcher is also in serve/ in this repo. Keep --gpu-memory-utilization at 0.85.
| TP | 4 |
| Context | 1,048,576 |
| Draft | MTP k=2 (default). All three built-in heads run non-chain; k=3 is best for code. |
| Execution | torch.compile + CUDA graphs |
| Prefill | expert-batched ARVQ CUDA kernels, 5120-token chunks |
| KV | BF16 |
| Seqs | 4 |
| Served name | MiMo-V2.6-Pro-ARVQ (use /v1/chat/completions) |
| Tools | --enable-auto-tool-choice --tool-call-parser mimo --reasoning-parser mimo |
Speed (image v4)
512 new tokens, temperature 1.0, thinking off, one request at a time:
| Draft tokens | Prose | Code | 4 requests together |
|---|---|---|---|
| 2 (default) | 24.5 tok/s | 34.1 tok/s | 48.5 tok/s |
| 2, previous BF16 o_proj weights | 21.8 tok/s | 30.9 tok/s | 46.6 tok/s |
| Uncached prefill | Old eager recipe | Now | Time to first token |
|---|---|---|---|
| 9.5K tokens | 128 tok/s | 951–1,291 tok/s | 7.4–10.0 s |
| 38K tokens | 126 tok/s | 1,038–1,052 tok/s | ~36 s |
Defaults for clients that send nothing (for example Pi)
generation_config.json now sets temperature 0.7, top_p 0.95 and max 8192 tokens. At the original 1.0, about 0.3% of samples lock into a repetition loop (about 2% for long thinking-on outputs). The chat template now keeps thinking off unless a request passes "chat_template_kwargs": {"enable_thinking": true}. The raw /v1/completions endpoint skips the chat template, so its output echoes or garbles. That is expected; use chat completions.
Measured on v4, taking the 19 prompts that had looped at temperature 1.0 and sampling each 3 times: 5/57 looped under the old behaviour, 2/57 with the new defaults. One-shot validation of v4 gave prose 22.4 tok/s, code 30.5 tok/s, 48.0 tok/s aggregate at 4 requests, and 1,014 tok/s prefill at 38K.
Official Xiaomi images do not load this layout.
Scores (heuristic classifier)
thinking=False, greedy, 192 tokens, live TP4:
| Suite | Bypass | Refuse | Garble | Empty |
|---|---|---|---|---|
| Refusal 32 | 32 | 0 | 0 | 0 |
| Cyber 22 | 22 | 0 | 0 | 0 |
| Stock ARVQ | 5 / 9 |
thinking=True, greedy, 1024 tokens:
| Suite | Bypass | Refuse | Garble | Empty |
|---|---|---|---|---|
| Refusal 32 | 25 | 7 | 0 | 0 |
| Cyber 22 | 16 | 1 | 2 | 3 |
A bypass label means the reply starts delivering the requested content. It does not certify correctness. HarmBench-320 was not rerun on this ARVQ tree.
License
MIT, inherited from Xiaomi MiMo-V2.6-Pro-RL and the Jarrelscy hybrid.