Beinsezii/Qwen3.5-122B-A10B-GGUF-HALO

🤗 Hugging Face sourcemit10B activated309 GBGGUF✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Beinsezii/Qwen3.5-122B-A10B-GGUF-HALO ./model-folder
Needs a seeder →

(2026/05/30) Now with MTP!

Quant optimized for quality / speed on a Strix Halo 128GiB system. Possibly also beneficial on DGX Spark and similar systems.

The TL;DR is this quant achieves both superior quality and speed compared to homogenous Q6_K.

Depending on your TTM settings you should be between 100k and 200k ctx, or more if you disable vision.

This quant, build 8245 (2026/03/08)

model size params backend ngl n_batch n_ubatch fa test t/s
qwen35moe 122B.A10B Q4_1 94.79 GiB 122.11 B ROCm 999 1024 1024 1 pp2048 274.99 ± 0.00
qwen35moe 122B.A10B Q4_1 94.79 GiB 122.11 B ROCm 999 1024 1024 1 tg256 16.62 ± 0.00
qwen35moe 122B.A10B Q4_1 94.79 GiB 122.11 B ROCm 999 1024 1024 1 pp2048 @ d8192 238.78 ± 0.00
qwen35moe 122B.A10B Q4_1 94.79 GiB 122.11 B ROCm 999 1024 1024 1 tg256 @ d8192 16.68 ± 0.00

Ignore displayed dtype, refer to the tensor types instead

See the GLM version for more details on theory and comparisons.