Qwythos-9B-v2-Heretic-MLX-4bit
MLX 4-bit quantization of WaveCut/Qwythos-9B-v2-Heretic — the Heretic-decensored version of empero-ai/Qwythos-9B-v2. Built for Apple Silicon (M1/M2/M3/M4).
Specs
| Field | Value |
|---|---|
| Bits/weight | 4.501 |
| File size | ~4.7 GB |
| Minimum RAM | ~6 GB unified memory |
Quantization
| Step | Tool | Version |
|---|---|---|
| Convert + quantize | mlx_lm.convert |
mlx-lm 0.31.3 (mlx 0.31.x) |
| Quant mode | affine (default) |
-q --q-bits 4 |
python -m mlx_lm.convert \
--hf-path WaveCut/Qwythos-9B-v2-Heretic \
-q --q-bits 4 \
--upload-repo WaveCut/Qwythos-9B-v2-Heretic-MLX-4bit
Usage
from mlx_lm import load, generate
model, tokenizer = load("WaveCut/Qwythos-9B-v2-Heretic-MLX-4bit")
response = generate(model, tokenizer, prompt="Hello", max_tokens=256)
print(response)
# CLI
mlx_lm.generate --model WaveCut/Qwythos-9B-v2-Heretic-MLX-4bit --prompt "Hello"
Architecture
Qwen3.5 hybrid — 32 blocks mixing attention and SSM (Mamba-style) layers. Supported in mlx-lm ≥ 0.31.0. Load with trust_remote_code=True only if your mlx-lm is older.
Disclaimer
Uncensored (safety alignment removed via Heretic). The original empero-ai/Qwythos-9B-v2 maintainers are not affiliated with this derivative. Use responsibly.