mlx-community/Qwen3.8-27B-Uncensored-OptiQ-4bit

🤗 Hugging Face 来源text-generationapache-2.026.9B 参数54 GBsafetensors✓ 6 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo mlx-community/Qwen3.8-27B-Uncensored-OptiQ-4bit ./model-folder
需要做种者 →

mlx-community/Qwen3.8-27B-Uncensored-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs

OptiQ mixed-precision quant of orcarouter/Qwen3.8-27B-Uncensored, a Qwen3.8-family reasoning model with a bundled MTP speculation head. 19 GB on disk.

What it is

Property Value
Base orcarouter/Qwen3.8-27B-Uncensored (Qwen3.8, 27B)
Method OptiQ mixed-precision, per-layer 4/8-bit
Bit allocation Reused from the Qwen3.8-27B OptiQ recipe: the architecture is identical, so the per-layer sensitivity ranking transfers directly and no per-model sweep is needed
Layer split 237 components at 4-bit, 261 at 8-bit
Group size 64
On disk 19 GB
MTP Speculation head preserved in optiq/mtp.safetensors for faster decode via optiq serve --draft-model

Following the naming llama.cpp uses for its mixed quants, the "4bit" label denotes the family, not the weighted average.

Run it

Qwen3.8 and the MTP sidecar register through OptiQ, so import optiq once before loading:

pip install "mlx-optiq>=0.4.27"
import optiq  # registers the arch + MTP sidecar
from mlx_lm import load, generate

model, tok = load("mlx-community/Qwen3.8-27B-Uncensored-OptiQ-4bit")
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Explain mixed-precision quantization in two sentences."}],
    tokenize=False, add_generation_prompt=True,
)
print(generate(model, tok, prompt=prompt, max_tokens=400))

For an OpenAI- and Anthropic-compatible endpoint with mixed-precision KV cache:

optiq serve --model mlx-community/Qwen3.8-27B-Uncensored-OptiQ-4bit

This is a reasoning model, so give it a generous token budget.

Links