Huihui-Qwen3.8-27B-abliterated — MLX MXFP4
MLX MXFP4 quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated,
uncensored via abliteration (refusal-direction removal on text layers 18–51; the README notes MTP and the vision tower were left unmodified) of Qwen3.8-27B. Qwen3.8-27B is a qwen3_5 vision-language model with
a hybrid GatedDeltaNet linear-attention + full-attention text backbone (64
layers, full attention every 4th) and a Qwen3-VL vision tower. Runs on Apple
Silicon via mlx-vlm. Stays
image-text-to-text — the vision tower is kept in bf16; only the text backbone
is quantized.
| Precision | MXFP4 (E2M1 + E8M0 shared scale, group size 32) |
| Bits per weight | 4.449 bpw |
| On-disk size | 14 GB (13 shards) |
| Quantized | text backbone (incl. the ~1.27B lm_head) |
| Kept in bf16 | Qwen3-VL vision tower |
Quantizations
| Variant | Bits | Size | |
|---|---|---|---|
Huihui-Qwen3.8-27B-abliterated-MXFP4 |
4.449 bpw | 14 GB | ← this repo |
Huihui-Qwen3.8-27B-abliterated-MXFP8 |
8.381 bpw | 27 GB | higher fidelity / for 32 GB+ |
Verification
Smoke-tested on Apple Silicon via mlx-vlm with deterministic greedy decoding, inspecting raw token IDs (not just detokenized text):
Text — coherent and correct:
- "What is the capital of France?" → "The capital of France is Paris."
- "What is 25 + 17?" → coherent step-by-step arithmetic.
Vision (the bf16 vision tower, through the quantized text backbone):
| Image | Question | Answer | |
|---|---|---|---|
| solid red | main color? | Red | ✅ |
| solid blue | main color? | Blue | ✅ |
| green circle | shape and color? | "green circle" | ✅ |
Color and shape are read correctly — the vision path is live.
Usage (mlx-vlm)
pip install -U mlx-vlm # needs the qwen3_5 architecture (>= 0.6.12)
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
model, processor = load("sahilchachra/Huihui-Qwen3.8-27B-abliterated-MXFP4")
config = model.config
prompt = apply_chat_template(processor, config, "What is the capital of France?")
print(generate(model, processor, prompt, max_tokens=256, verbose=True))
This is a reasoning model; it produces a <think> channel before its answer.
Run in LM Studio
Loads and runs in LM Studio (tested on 0.4.20, mlx-llm runtime): the
qwen3_5 architecture is recognized, the model indexes cleanly (format: mlx
present), and the ChatML template runs as-is. This is a reasoning model — it
emits a thinking channel (reasoning_content) before the final content, so give
it enough max_tokens (e.g. 200+) or the answer can be empty while it is still
thinking. Verified: "capital of France" → reasoning + content = "Paris".
Loads on a 24 GB Mac (14.2 GiB weights + small context).
Notes & limitations
- Uncensored model. This is a deliberately uncensored/abliterated derivative and will not refuse requests the way the original might. Use responsibly and in line with the base model's license and your local laws.
- MTP head dropped. The base model's multi-token-prediction (speculative
decoding) head is not included — mlx-vlm strips the
mtp.*weights on load. Text and vision are unaffected; only spec-decode is not available. tie_word_embeddings=false, so the ~1.27Blm_headis a separate matrix and is quantized. Verified correct on the MXFP4 build above.- Inherits all capabilities and limitations of the base model. See the original model card.
- Quantized by @sahilchachra with MLX. Apache-2.0.