npario/Qwen3.8-27B-OBLITERATED-OptiQ-4bit

🤗 Hugging Face sourceimage-text-to-textapache-2.026.9B params55 GBsafetensors✓ 7 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo npario/Qwen3.8-27B-OBLITERATED-OptiQ-4bit ./model-folder
Needs a seeder →

mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs

OptiQ mixed-precision quant of OBLITERATUS/Qwen3.8-27B-OBLITERATED, a Qwen3.8-family vision-language model with a bundled MTP speculation head. 21 GB on disk.

What it is

Property Value
Base OBLITERATUS/Qwen3.8-27B-OBLITERATED (Qwen3.8, 27B)
Method OptiQ mixed-precision, per-layer 4/8-bit
Bit allocation Reused from the Qwen3.8-27B OptiQ recipe: the architecture is identical, so the per-layer sensitivity ranking transfers directly and no per-model sweep is needed
Layer split 237 components at 4-bit, 261 at 8-bit
Group size 64
On disk 21 GB
MTP Speculation head preserved in optiq/mtp.safetensors
Vision bf16 vision tower kept in optiq/optiq_vision.safetensors for image input

Following the naming llama.cpp uses for its mixed quants, the "4bit" label denotes the family, not the weighted average.

Run it

Qwen3.8 and the MTP/vision sidecars register through OptiQ, so import optiq once before loading:

pip install "mlx-optiq>=0.4.27"
import optiq  # registers the arch + MTP/vision sidecars
from mlx_lm import load, generate

model, tok = load("mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit")
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Explain mixed-precision quantization in two sentences."}],
    tokenize=False, add_generation_prompt=True,
)
print(generate(model, tok, prompt=prompt, max_tokens=400))

For image input plus an OpenAI- and Anthropic-compatible endpoint with mixed-precision KV cache:

optiq serve --model mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit

This is a reasoning model, so give it a generous token budget.

Links