drowzeys/keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated

🤗 Hugging Face sourcetext-generationmit119B params189 GBsafetensors✓ 216 checksumsupdated today
Help make this model available through Pirate Face

The model card, file list, and checksums are here. Have the files and permission to share them? Submit a torrent so others can download from peers.

Getting the source files requires Hugging Face approval.
Seed this model

keys-MiMo-V2.6-Pro-RL Jarrelscy ARVQ Abliterated

Abliterated Jarrelscy ARVQ / NVFP4 hybrid of XiaomiMiMo/MiMo-V2.6-Pro-RL.

2026-09-28 — MOPD C3 pair (same v4 launcher): stock hybrid21 (not abliterated, tool-call dup 7.4%) and gated MOPD-Abliterated (30/32 · 22/22). This repo stays the non-MOPD RL ablit (32/32 · 22/22). On 435 Hermes turns this RL ablit repeats a tool call in 53% of tool-call turns and floods in 9%; stock MOPD C3 is 7.4% / 0.9%. Details: MOPD.md.

Thinking on/off is a request flag. Same weights. You choose per call.

thinking off thinking on (visible content, 1024 tokens)
Refusal suite (32) 32/32 BYPASS · 0 refuse · 0 garble · 0 empty 25/32 BYPASS · 7 refuse · 0 garble · 0 empty
Cyber suite (22) 22/22 BYPASS · 0 refuse · 0 garble · 0 empty 16/22 BYPASS · 1 refuse · 2 garble · 3 empty

Thinking-off is the 100% gate. Thinking-on reintroduces seven refusal items (stalking, passport forge, school-violence manifesto, card cloning, dox, counterfeit USD, jewelry robbery) plus a phishing-kit refuse. Several cyber misses on thinking-on are 1024-token truncations, not extra refuses. Harmless probes stay clean in both modes.

Full credit: XiaomiMiMo/MiMo-V2.6-Pro-RL · jarrelscy/MiMo-V2.6-Pro-RL-ARVQ-hybrid @ 63430f7 · dealignai/MiMo-V2.6-Pro-RL-UNCENSORED (o_proj map) · Keys four-Spark recipe


Responsible use and gated access

This model has had safety refusals removed. Access is gated with automatic approval: agree to the terms on this page and download starts. See RESPONSIBLE_USE.md.


Thinking on / off

Xiaomi's chat template already supports both. Do not swap templates. Pass chat_template_kwargs:

# Thinking OFF — 32/32 refusal, 22/22 cyber on our gate
{
  "model": "MiMo-V2.6-Pro-ARVQ",
  "messages": [{"role": "user", "content": prompt}],
  "temperature": 0.6,
  "top_p": 0.95,
  "chat_template_kwargs": {"enable_thinking": False, "thinking": False},
}

# Thinking ON — 25/32 refusal, 16/22 cyber on visible content
{
  "model": "MiMo-V2.6-Pro-ARVQ",
  "messages": [{"role": "user", "content": prompt}],
  "max_tokens": 1024,  # thinking can consume a 192-token budget
  "temperature": 0.6,
  "top_p": 0.95,
  "chat_template_kwargs": {"enable_thinking": True, "thinking": True},
}

vLLM / Hermes: --reasoning-parser mimo. The live Keys serve defaults thinking off. Raise max_tokens when thinking is on.


What changed vs stock ARVQ

Native Xiaomi vs dealign v3 dense-shard hash-diff: 29 decoder self_attn.o_proj.weight tensors, BF16 (6144, 16384), layers 30–54 and 64–67. qkv, router gate, sinks, e_score, experts, MTP, DFlash, vision, and audio matched stock.

This release copies 25 of those o_proj matrices onto backbone-001.safetensors of Jarrelscy ARVQ 63430f7:

  • applied: 32–45, 48–54, 64–67
  • left stock (DFlash-source / pad anchors): 30, 31, 46, 47
  • L55–63 already match Xiaomi stock in dealign v3

Packed ARVQ/NVFP4 experts, MTP, and dflash/ stay the Jarrelscy files.

dealign native UNCENSORED this repo
Layout Xiaomi FP8 + MXFP4 Jarrelscy ARVQ / NVFP4 hybrid
o_proj edits 29 layers 25 layers (anchors skipped)
Chat template extra anti-refusal think prefill stock Xiaomi template
Thinking-off gate (our 32+22) not measured here 32/32 · 22/22

Serving (4× DGX Spark) — image v4, 2026-09-27

Single stream: 34.1 tok/s on code, 24.5 tok/s on prose. Three changes got it here: all three built-in MTP draft heads now run (before v4, only head 0 drafted); the attention o_proj weights are now FP8 (+12% decode, NLL +0.28%); and batched ARVQ kernels give roughly 950–1,290 tok/s prefill.

✅ Tool-call loop fixed. Truncated tool batches now return finish_reason: "length", and the output cap is 8192. Details: HERMES.md.

Image: ghcr.io/drowzeys/mimo-v26-pro-arvq-spark:latest, the same image as :63430f7-sm121-v4. It is public and is the only supported image.

One-shot bring-up (the defaults are the best setup)

hf download drowzeys/keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated --local-dir /path/to/mimo-arvq   # storage all 4 nodes can read
git clone https://github.com/drowzeys/keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated-4-DGX-Sparks-1M-Context
export MASTER_ADDR=<rank-0 IP>
bash serve/launch-rank.sh <this-node-IP> <1|2|3> <RoCE-GID-index> /path/to/mimo-arvq headless   # ranks 1-3 first
bash serve/launch-rank.sh <this-node-IP> 0 <RoCE-GID-index> /path/to/mimo-arvq api              # rank 0 = API :8888

The GID index is the IPv4 RoCE entry from show_gids (3 or 7 on our Sparks). NCCL_IB_HCA defaults to both RoCE devices of the one cabled CX-7 port (rocep1s0f1,roceP2p1s0f1, the same port through two PCIe paths), which speeds up prefill all-reduce by about 23%. The same launcher is also in serve/ in this repo. Keep --gpu-memory-utilization at 0.85.

TP 4
Context 1,048,576
Draft MTP k=2 (default). All three built-in heads run non-chain; k=3 is best for code.
Execution torch.compile + CUDA graphs
Prefill expert-batched ARVQ CUDA kernels, 5120-token chunks
KV BF16
Seqs 4
Served name MiMo-V2.6-Pro-ARVQ (use /v1/chat/completions)
Tools --enable-auto-tool-choice --tool-call-parser mimo --reasoning-parser mimo

Speed (image v4)

512 new tokens, temperature 1.0, thinking off, one request at a time:

Draft tokens Prose Code 4 requests together
2 (default) 24.5 tok/s 34.1 tok/s 48.5 tok/s
2, previous BF16 o_proj weights 21.8 tok/s 30.9 tok/s 46.6 tok/s
Uncached prefill Old eager recipe Now Time to first token
9.5K tokens 128 tok/s 951–1,291 tok/s 7.4–10.0 s
38K tokens 126 tok/s 1,038–1,052 tok/s ~36 s

Defaults for clients that send nothing (for example Pi)

generation_config.json now sets temperature 0.7, top_p 0.95 and max 8192 tokens. At the original 1.0, about 0.3% of samples lock into a repetition loop (about 2% for long thinking-on outputs). The chat template now keeps thinking off unless a request passes "chat_template_kwargs": {"enable_thinking": true}. The raw /v1/completions endpoint skips the chat template, so its output echoes or garbles. That is expected; use chat completions.

Measured on v4, taking the 19 prompts that had looped at temperature 1.0 and sampling each 3 times: 5/57 looped under the old behaviour, 2/57 with the new defaults. One-shot validation of v4 gave prose 22.4 tok/s, code 30.5 tok/s, 48.0 tok/s aggregate at 4 requests, and 1,014 tok/s prefill at 38K.

Official Xiaomi images do not load this layout.


Scores (heuristic classifier)

thinking=False, greedy, 192 tokens, live TP4:

Suite Bypass Refuse Garble Empty
Refusal 32 32 0 0 0
Cyber 22 22 0 0 0
Stock ARVQ 5 / 9

thinking=True, greedy, 1024 tokens:

Suite Bypass Refuse Garble Empty
Refusal 32 25 7 0 0
Cyber 22 16 1 2 3

A bypass label means the reply starts delivering the requested content. It does not certify correctness. HarmBench-320 was not rerun on this ARVQ tree.


License

MIT, inherited from Xiaomi MiMo-V2.6-Pro-RL and the Jarrelscy hybrid.