WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit

🤗 Hugging Face sourcetext-generationapache-2.09B params18 GBsafetensors✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit ./model-folder
Needs a seeder →

Qwythos-9B-v2-Heretic-MLX-8bit

MLX 8-bit quantization of WaveCut/Qwythos-9B-v2-Heretic — the Heretic-decensored version of empero-ai/Qwythos-9B-v2. Built for Apple Silicon (M1/M2/M3/M4).

Specs

Field Value
Bits/weight 8.501
File size ~9.5 GB
Minimum RAM ~11 GB unified memory

Quantization

Step Tool Version
Convert + quantize mlx_lm.convert mlx-lm 0.31.3 (mlx 0.31.x)
Quant mode affine (default) -q --q-bits 8
python -m mlx_lm.convert \
  --hf-path WaveCut/Qwythos-9B-v2-Heretic \
  -q --q-bits 8 \
  --upload-repo WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit

Usage

from mlx_lm import load, generate
model, tokenizer = load("WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit")
response = generate(model, tokenizer, prompt="Hello", max_tokens=256)
print(response)
# CLI
mlx_lm.generate --model WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit --prompt "Hello"

Architecture

Qwen3.5 hybrid — 32 blocks mixing attention and SSM (Mamba-style) layers. Supported in mlx-lm ≥ 0.31.0.

Disclaimer

Uncensored (safety alignment removed via Heretic). The original empero-ai/Qwythos-9B-v2 maintainers are not affiliated with this derivative. Use responsibly.