Beinsezii/Qwen3.5-122B-A10B-GGUF-HALO

🤗 Hugging Face 来源mit激活 10B309 GBGGUF✓ 3 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Beinsezii/Qwen3.5-122B-A10B-GGUF-HALO ./model-folder
需要做种者 →

(2026/05/30) Now with MTP!

Quant optimized for quality / speed on a Strix Halo 128GiB system. Possibly also beneficial on DGX Spark and similar systems.

The TL;DR is this quant achieves both superior quality and speed compared to homogenous Q6_K.

Depending on your TTM settings you should be between 100k and 200k ctx, or more if you disable vision.

This quant, build 8245 (2026/03/08)

model size params backend ngl n_batch n_ubatch fa test t/s
qwen35moe 122B.A10B Q4_1 94.79 GiB 122.11 B ROCm 999 1024 1024 1 pp2048 274.99 ± 0.00
qwen35moe 122B.A10B Q4_1 94.79 GiB 122.11 B ROCm 999 1024 1024 1 tg256 16.62 ± 0.00
qwen35moe 122B.A10B Q4_1 94.79 GiB 122.11 B ROCm 999 1024 1024 1 pp2048 @ d8192 238.78 ± 0.00
qwen35moe 122B.A10B Q4_1 94.79 GiB 122.11 B ROCm 999 1024 1024 1 tg256 @ d8192 16.68 ± 0.00

Ignore displayed dtype, refer to the tensor types instead

See the GLM version for more details on theory and comparisons.