Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-mxfp4-mlx
Brainwaves
arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.665,0.831,0.910,0.790,0.456,0.813,0.772
Quant Perplexity Peak Memory Tokens/sec
mxfp8 5.073 ± 0.038 34.74 GB 212
mxfp4 4.950 ± 0.036 21.30 GB 187
Baseline model
arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.647,0.803,0.910,0.773,0.450,0.806,0.742
qx86-hi 0.637,0.798,0.911,0.775,0.442,0.807,0.737
This model is using the fixed jinja template from froggeric/Qwen-Fixed-Chat-Templates
Thinking toggle
Drop <|think_on|> or <|think_off|> anywhere in your system or user prompt. The template intercepts the tag, removes it from context so the model never sees it, and flips the mode.
Fast answer, no reasoning:
System: You are a coding assistant. <|think_off|>
User: What's 2+2?
Deep reasoning:
System: You are a coding assistant. <|think_on|>
User: Implement a red-black tree in Rust.
The tag syntax (<|think_on|>, <|think_off|>) uses Qwen's control-token delimiters, so it will never collide with real text. Earlier community templates used /think, which broke legitimate paths like cd /mnt/project/think.
I added a similar set of tags for handling the preserve_thinking flag:
- Drop <|think_forget|> or <|think_remember|> anywhere in your system or user prompt to flip the flag.
- The template intercepts the tag, removes it from context so the model never sees it, and flips the mode.
-G
Use with mlx
pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-mxfp4-mlx")
prompt = "hello"
if tokenizer.chat_template is not None:
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_dict=False,
)
response = generate(model, tokenizer, prompt=prompt, verbose=True)