Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-1M-qx86-hi-mlx
This model is a merge of:
- armand0e/Qwen3.6-35B-A3B-Fable-5-Distill
- Hcompany/Holo3.1-35B-A3B
- Jackrong/Qwopus3.6-35B-A3B-Coder
Let me know if it worked for you.
If this model gets more Likes, I will provide the source--usually not here because of space constraints.
Brainwaves
arc arc/e boolq hswag obkqa piqa wino
bf16 0.645,0.837,0.894,0.783,0.454,0.822,0.735
mxfp8 0.645,0.833,0.894,0.783,0.454,0.820,0.725
qx86-hi 0.647,0.843,0.893,0.780,0.446,0.822,0.730
qx64-hi 0.655,0.839,0.894,0.778,0.442,0.824,0.725
mxfp4 0.637,0.832,0.889,0.776,0.462,0.817,0.714
Similar model in this range
Jiunsong/SuperQwen-AgentWorld-35B-A3B-abliterated
arc arc/e boolq hswag obkqa piqa wino
mxfp4 0.657,0.862,0.906,0.766,0.490,0.825,0.692
Qwen/Qwen-AgentWorld-35B-A3B
arc arc/e boolq hswag obkqa piqa wino
qx64-hi 0.644,0.818,0.909
mxfp4 0.626,0.813,0.901
Model components
armand0e/Qwen3.6-35B-A3B-Fable-5-Distill
arc arc/e boolq hswag obkqa piqa wino
qx86-hi 0.635,0.821,0.891,0.770,0.444,0.818,0.721
Hcompany/Holo-3.1-35B-A3B
arc arc/e boolq hswag obkqa piqa wino
qx86-hi 0.533,0.705,0.882,0.771,0.456,0.811,0.690
Jackrong/Qwopus3.6-35B-A3B-Coder
arc arc/e boolq hswag obkqa piqa wino
qx86-hi 0.594,0.770,0.888,0.750,0.438,0.813,0.717
Baseline model
Qwen3.6-35B-A3B-Instruct
arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.581,0.757,0.892,0.751,0.428,0.803,0.688
qx86-hi 0.576,0.742,0.896,0.745,0.422,0.803,0.708
mxfp4 0.586,0.767,0.886,0.751,0.428,0.798,0.681
Quant Perplexity Peak Memory Tokens/sec
mxfp8 5.138 ± 0.037 42.65 GB 1201
mxfp4 5.158 ± 0.037 25.33 GB 1355
qx86-hi 4.826 ± 0.033 45.50 GB 1474
qx64-hi 4.710 ± 0.032 36.83 GB 1414
Thinking toggle
This model is using the fixed jinja template from froggeric/Qwen-Fixed-Chat-Templates
Contribute to NightmediaAI
If you like our models and want to contribute to help us improve our lab, any form would do:
ETH:0x6b6633606995BC180925c47d4249ED624aB7b2A5 USDC:0x19e6bDDCBa47BB09a9Bc153Bb6479fc57284421a BTC:36d7U1n3MFaXgnNRAaEL3Pa3Hy6oFhM7XY BCH:15dNMzhJ87XJSTU89VCBsDHj747QvBQaap
My models and I thank you :)
-G
Use with mlx
pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-1M-qx86-hi-mlx")
prompt = "hello"
if tokenizer.chat_template is not None:
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_dict=False,
)
response = generate(model, tokenizer, prompt=prompt, verbose=True)