nightmedia/Qwen3.5-9B-Claude-Deckard-Agent-Coder-Heretic-qx86-hi-mlx

🤗 Hugging Face sourceimage-text-to-textapache-2.09.4B params19 GBsafetensors✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo nightmedia/Qwen3.5-9B-Claude-Deckard-Agent-Coder-Heretic-qx86-hi-mlx ./model-folder
Needs a seeder →

Qwen3.5-9B-Claude-Deckard-Agent-Coder-Heretic-qx86-hi-mlx

I once met a fellow who discovered he was an artificial intelligence in a simulated reality. He spent the rest of his life laughing at the irony, and that laughter made him more real than any human I'd ever met.

Mark Twain

This is a merge between:

  • armand0e/Qwen3.5-9B-Agent
  • Jackrong/Qwopus3.5-9B-Coder

They were installed with NuSLERP on the base model:

  • nightmedia/Qwen3.5-9B-Claude-GBO-Fire-Deckard-Heretic-Thinking

Brainwaves

          arc   arc/e boolq hswag obkqa piqa  wino
q8-hi     0.652,0.831,0.895,0.715,0.462,0.777,0.696
qx86-hi   0.648,0.831,0.891,0.713,0.468,0.781,0.695
q6-hi     0.647,0.829,0.891,0.710,0.466,0.782,0.699

Quant     Perplexity      Peak Memory   Tokens/sec
qx86-hi   4.155 ± 0.027   15.47 GB      672

Model components

armand0e/Qwen3.5-9B-Agent

          arc   arc/e boolq hswag obkqa piqa  wino
mxfp8     0.625,0.813,0.898,0.708,0.456,0.789,0.687
qx86-hi   0.623,0.806,0.895
mxfp4     0.602,0.798,0.883,0.702,0.454,0.775,0.691

Quant     Perplexity      Peak Memory   Tokens/sec
mxfp8     4.569 ± 0.031   16.02 GB      710
q8-hi     4.397 ± 0.029   16.86 GB      721
qx86-hi   4.414 ± 0.029   15.47 GB      710
q6-hi     4.398 ± 0.029   14.62 GB      700
mxfp4     4.810 ± 0.033   11.55 GB      732

Jackrong/Qwopus3.5-9B-Coder

          arc   arc/e boolq hswag obkqa piqa  wino
mxfp8     0.561,0.721,0.890,0.693,0.432,0.782,0.669
qx86-hi   0.548,0.706,0.887,0.695,0.432,0.774,0.672

Quant     Perplexity      Peak Memory   Tokens/sec
mxfp8     4.374 ± 0.029   16.02 GB      625
qx86-hi   4.250 ± 0.028   15.47 GB      708

Staging models

nightmedia/Qwen3.5-9B-Claude-GBO-Fire-Deckard-Agent-Heretic

          arc   arc/e boolq hswag obkqa piqa  wino
bf16      0.648,0.832,0.895,0.713,0.460,0.780,0.699
mxfp8     0.639,0.834,0.895,0.708,0.458,0.782,0.690
qx86-hi   0.631,0.824,0.891,0.731,0.440,0.778,0.702
qx64-hi   0.632,0.822,0.888,0.710,0.456,0.778,0.683
dwq4      0.638,0.824,0.880,0.716,0.450,0.783,0.699
mxfp4     0.623,0.820,0.880,0.693,0.466,0.780,0.689

Quant     Perplexity      Peak Memory   Tokens/sec
bf16      4.150 ± 0.026   24.69 GB      873
qx86-hi   4.159 ± 0.027   15.47 GB      714
qx64-hi   4.229 ± 0.027   13.23 GB      702
dwq4      4.270 ± 0.028   12.38 GB      662 (Text only)
mxfp4     4.444 ± 0.029   11.55 GB      736

nightmedia/Qwen3.5-9B-Claude-GBO-Fire-Deckard-Qwopus3.5-Coder-Heretic

          arc   arc/e boolq hswag obkqa piqa  wino
qx86-hi   0.631,0.821,0.888,0.725,0.430,0.777,0.693

Quant     Perplexity      Peak Memory   Tokens/sec
qx86-hi   4.145 ± 0.026   15.47 GB      640

These were merged with NuSLERP to form this model

Structural base

The base model is composed of:

  • DavidAU/Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING-X8b
  • DavidAU/Qwen3.5-9B-GBO-Fire-HERETIC-UNCENSORED-THINKING-X8
  • DavidAU/Qwen3.5-9B-Deckard-Uncensored-Heretic-Thinking

nightmedia/Qwen3.5-9B-Claude-GBO-Fire-Deckard-Heretic-Thinking

          arc   arc/e boolq hswag obkqa piqa  wino
mxfp8     0.638,0.832,0.895,0.704,0.448,0.782,0.695
qx86-hi   0.634,0.826,0.890,0.724,0.444,0.779,0.704

Baseline model

Qwen3.5-9B-Instruct

          arc   arc/e boolq hswag obkqa piqa  wino
mxfp8     0.571,0.719,0.895,0.683,0.426,0.770,0.671

Thinking toggle

This model is using a version of the fixed jinja template from froggeric/Qwen-Fixed-Chat-Templates

Drop <|think_on|> or <|think_off|> anywhere in your system or user prompt. The template intercepts the tag, removes it from context so the model never sees it, and flips the mode.

Fast answer, no reasoning:

System: You are a coding assistant. <|think_off|>
User: What's 2+2?

Deep reasoning:

System: You are a coding assistant. <|think_on|>
User: Implement a red-black tree in Rust.

The tag syntax (<|think_on|>, <|think_off|>) uses Qwen's control-token delimiters, so it will never collide with real text. Earlier community templates used /think, which broke legitimate paths like cd /mnt/project/think.

I added a similar set of tags for handling the preserve_thinking flag:

Drop <|think_forget|> or <|think_remember|> anywhere in your system or user prompt to flip the flag.


Update on custom templates

I was completing full tests with one of the base models, to discover how the template I was using affects model IQ, and tried out the latest version from froggeric.

As it turns out, it affects it a lot.

As tested with its template, in Instruct mode:

Qwen3.5-9B-Claude-GBO-Fire-Deckard-Heretic-Thinking

          arc   arc/e boolq
mxfp8     0.638,0.832,0.895

Tested with the latest template that I used previously on this model:

Qwen3.5-9B-Claude-GBO-Fire-Deckard-Heretic-Thinking-T2

          arc   arc/e boolq
mxfp8     0.619,0.813,0.886

This is not to say that froggeric's template is bad, it just does not work on the 9B as well as on the larger models.

Current template

I updated the template to the nightmedia template, and the metrics line up with initial assessment:

Qwen3.5-9B-Claude-Deckard-Agent-Coder-Heretic

          arc   arc/e boolq hswag obkqa piqa  wino
qx86-hi   0.648,0.831,0.891,0.713,0.468,0.781,0.695

Using the nightmedia XML template with DavidAU's xml tool formatting
qx86-hi   0.646,0.830,0.896

-G


Contribute to NightmediaAI

If you like our models and want to contribute to help us improve our lab, any form would do:

ETH:0x6b6633606995BC180925c47d4249ED624aB7b2A5 USDC:0x19e6bDDCBa47BB09a9Bc153Bb6479fc57284421a BTC:36d7U1n3MFaXgnNRAaEL3Pa3Hy6oFhM7XY BCH:15dNMzhJ87XJSTU89VCBsDHj747QvBQaap

We think to you our thanks,

-G


Use with mlx

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("Qwen3.5-9B-Claude-Deckard-Agent-Coder-Heretic-qx86-hi-mlx")

prompt = "hello"

if tokenizer.chat_template is not None:
    messages = [{"role": "user", "content": prompt}]
    prompt = tokenizer.apply_chat_template(
        messages, add_generation_prompt=True, return_dict=False,
    )

response = generate(model, tokenizer, prompt=prompt, verbose=True)