boboliu/Qwen3-Embedding-8B-W4A16-G128

🤗 Hugging Face sourcefeature-extractionapache-2.08.2B params30 GBsafetensors✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo boboliu/Qwen3-Embedding-8B-W4A16-G128 ./model-folder
Needs a seeder →

Qwen3-Embedding-8B-W4A16-G128

GPTQ Quantized https://huggingface.co/Qwen/Qwen3-Embedding-8B with THUIR/T2Ranking and m-a-p/COIG-CQIA for calibration set.

What's the benefit?

VRAM Usage: more than 24G -> 19624M, make it available on 3090/4090. (w/o FA2)

What's the cost?

~0.81% lost in C-MTEB.

C-MTEB Param. Mean(Task) Mean(Type) Class. Clust. Pair Class. Rerank. Retr. STS
multilingual-e5-large-instruct 0.6B 58.08 58.24 69.80 48.23 64.52 57.45 63.65 45.81
bge-multilingual-gemma2 9B 67.64 75.31 59.30 86.67 68.28 73.73 55.19 -
gte-Qwen2-1.5B-instruct 1.5B 67.12 67.79 72.53 54.61 79.5 68.21 71.86 60.05
gte-Qwen2-7B-instruct 7.6B 71.62 72.19 75.77 66.06 81.16 69.24 75.70 65.20
ritrieve_zh_v1 0.3B 72.71 73.85 76.88 66.5 85.98 72.86 76.97 63.92
Qwen3-Embedding-8B 8B 73.84 75.00 76.97 80.08 84.23 66.99 78.21 63.53
This Model 8B-W4A16 73.24 74.38 76.85 79.58 83.21 66.43 77.39 62.80

How to use it?

pip install compressed-tensors optimum and auto-gptq / gptqmodel, then goto the official usage guide.