- Q4F : Q4_K feed-forawrd (Q5_1 for ffn_down due to shape constraints)
- Q8A : Q8_0 attention, Q8_0 output, Q8_0 embeds
- Q8SH : Q8_0 shared experts
Readable speeds on a 24GiB GPU + 64GB RAM
在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。
curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Beinsezii/GLM-4.6V-Q4F-Q8A-Q8SH-GGUF ./model-folderReadable speeds on a 24GiB GPU + 64GB RAM