Qwythos-9B-v2-Heretic-MLX-8bit
MLX 8-bit quantization of WaveCut/Qwythos-9B-v2-Heretic — the Heretic-decensored version of empero-ai/Qwythos-9B-v2. Built for Apple Silicon (M1/M2/M3/M4).
Specs
| Field | Value |
|---|---|
| Bits/weight | 8.501 |
| File size | ~9.5 GB |
| Minimum RAM | ~11 GB unified memory |
Quantization
| Step | Tool | Version |
|---|---|---|
| Convert + quantize | mlx_lm.convert |
mlx-lm 0.31.3 (mlx 0.31.x) |
| Quant mode | affine (default) |
-q --q-bits 8 |
python -m mlx_lm.convert \
--hf-path WaveCut/Qwythos-9B-v2-Heretic \
-q --q-bits 8 \
--upload-repo WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit
Usage
from mlx_lm import load, generate
model, tokenizer = load("WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit")
response = generate(model, tokenizer, prompt="Hello", max_tokens=256)
print(response)
# CLI
mlx_lm.generate --model WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit --prompt "Hello"
Architecture
Qwen3.5 hybrid — 32 blocks mixing attention and SSM (Mamba-style) layers. Supported in mlx-lm ≥ 0.31.0.
Disclaimer
Uncensored (safety alignment removed via Heretic). The original empero-ai/Qwythos-9B-v2 maintainers are not affiliated with this derivative. Use responsibly.