npario/Qwen3.8-27B-OBLITERATED-OptiQ-4bit

🤗 Hugging Face 来源image-text-to-textapache-2.026.9B 参数55 GBsafetensors✓ 7 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo npario/Qwen3.8-27B-OBLITERATED-OptiQ-4bit ./model-folder
需要做种者 →

mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs

OptiQ mixed-precision quant of OBLITERATUS/Qwen3.8-27B-OBLITERATED, a Qwen3.8-family vision-language model with a bundled MTP speculation head. 21 GB on disk.

What it is

Property Value
Base OBLITERATUS/Qwen3.8-27B-OBLITERATED (Qwen3.8, 27B)
Method OptiQ mixed-precision, per-layer 4/8-bit
Bit allocation Reused from the Qwen3.8-27B OptiQ recipe: the architecture is identical, so the per-layer sensitivity ranking transfers directly and no per-model sweep is needed
Layer split 237 components at 4-bit, 261 at 8-bit
Group size 64
On disk 21 GB
MTP Speculation head preserved in optiq/mtp.safetensors
Vision bf16 vision tower kept in optiq/optiq_vision.safetensors for image input

Following the naming llama.cpp uses for its mixed quants, the "4bit" label denotes the family, not the weighted average.

Run it

Qwen3.8 and the MTP/vision sidecars register through OptiQ, so import optiq once before loading:

pip install "mlx-optiq>=0.4.27"
import optiq  # registers the arch + MTP/vision sidecars
from mlx_lm import load, generate

model, tok = load("mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit")
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Explain mixed-precision quantization in two sentences."}],
    tokenize=False, add_generation_prompt=True,
)
print(generate(model, tok, prompt=prompt, max_tokens=400))

For image input plus an OpenAI- and Anthropic-compatible endpoint with mixed-precision KV cache:

optiq serve --model mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit

This is a reasoning model, so give it a generous token budget.

Links