agentionai/Qwen3.8-27B-AP-GGUF

🤗 Hugging Face 来源image-text-to-textapache-2.0激活 27B196 GBGGUF✓ 15 个校验和今天更新
已有模型文件?提交模型种子

如果你有完整的模型文件并有权分享,请把示例文件夹路径替换为你的文件路径,再运行这条命令。它会校验文件、制作种子,并将磁力链接和校验和提交给 Pirate Face。请让种子客户端持续做种,方便其他人从节点下载。Pirate Face 不接收模型文件。你可以从账户页面获取社区密钥。也可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo agentionai/Qwen3.8-27B-AP-GGUF ./model-folder
需要做种者 →

Qwen3.8-27B · Agention Precision GGUF

Same size. Same speed. Closer to the model Qwen trained than any other quant

Full Results

Agention Precision quants are the highest precision quants of Qwen3.8 27B byte-for-byte. Each file has a non-uniform assignment and low-loss error correction built with a custom encoder.

This is a drop-in GGUF pack for Qwen3.8-27B: standard llama.cpp types, no fork, no flags. Vision projector and MTP draft head included.

Every 27B quant gives something up. These give up less. We measured the leading public GGUFs against the full BF16 model on one protocol. At every size we ship, ours is the closest to Qwen’s next-token distribution.

Every quant is tested against three corpora: unseen technical prose, general web text and Wikipedia.

Swap cost is zero. VRAM and tokens/sec stay the same. The file behaves more like the weights Qwen released.

Start here File Size 32k VRAM Fits Gain vs leading same-size quant
Most headroom AP-Q4_K_XL 16.35 GiB ~18 GiB 24 GB 4% closer to BF16; hardest 1% of tokens 7% closer
24 GB, smaller AP-Q4_K_M 15.33 GiB ~17 GiB 24 GB 5% closer on unseen technical text; hardest 1% of tokens 9% closer
Default AP-IQ4_XS 13.27 GiB ~15 GiB 16 GB 8% closer to BF16; hardest 1% of tokens 9% closer
Need headroom AP-Q3_K_XL 12.24 GiB ~14 GiB 16 GB + longer ctx 19% closer on unseen technical text, 6% on web; hardest 1% of tokens 23% closer
12 GB, more quality AP-IQ3_S 11.21 GiB ~13 GiB 16 GB at 8–16k 17–20% closer on unseen technical text than the leading research quants, 5% on web; hardest 1% of tokens 24% closer
12 GB, balanced AP-IQ3_XS 10.70 GiB ~12.5 GiB 12 GB at 8–16k Level with ISTA’s 11.29 GiB IQ3_S on unseen technical text at 0.6 GiB less; 49% closer than the 10.38 GiB AtomicChat quant
12 GB / multi-model AP-IQ3_XXS 10.00 GiB ~12 GiB 12 GB at 8–16k 26% closer on unseen technical text, 18% on web and wikitext-2 than the best ~10 GiB research quant
12 GB, smallest AP-IQ2_S 8.95 GiB ~10.9 GiB 12 GB at 8–16k 22% closer on unseen technical text, 14% on web, 18% on wikitext-2 than the best same-size research quant
Vision mmproj-BF16.gguf 0.87 GiB +0.9 GiB any tier Qwen’s own encoder at BF16

VRAM = weights + q8_0 KV + llama.cpp buffers. Only a quarter of layers use full attention, so 32k context is about 1 GiB at q8_0.

Take AP-IQ4_XS if it fits. Step down only for memory.


Measured

Fidelity = KL divergence of the next-token distribution vs Qwen3.8-27B BF16. Lower is closer to the original.

60 × 2048 tokens, three corpora, same build, same BF16 logits:

  • held-out — 1.85 MB of our own technical prose, never part of any calibration set. It is internal engineering documentation, so it is not published: this column is the one you cannot re-run yourself.
  • neutral web — mixedweb-v1, 800,789 chars, a seeded FineWeb slice, md5 51e0045e8cabf37922aa82766a25b7b4, published with its per-document manifest and builder in agentionai/quant-fidelity-corpora
  • wikitext-2 — the standard wiki.test.raw, 1,288,556 chars, md5 7c0137fc034ddbc56a296bce31b4f7fb (llama.cpp/scripts/get-wikitext-2.sh)

Reproducing these numbers

The reference is Qwen/Qwen3.8-27B at revision 1d4bf0f2 (the weights uploaded 2026-08-13; later commits are README only), converted with llama.cpp's own converter — no changes, no fork:

python convert_hf_to_gguf.py --outtype bf16 Qwen3.8-27B/ --outfile Qwen3.8-27B-BF16.gguf   # 54,657,733,888 bytes

Then one base of BF16 logits per corpus, and every quant scored against it. The KLD code is upstream and unmodified; ours was built at llama.cpp 26bc85e42:

llama-perplexity -m Qwen3.8-27B-BF16.gguf -f mixedweb-v1.txt \
  --kl-divergence-base base-mixedweb.bin -c 2048 --chunks 60 -ngl 99

llama-perplexity -m Qwen3.8-27B-AP-IQ4_XS.gguf --kl-divergence \
  --kl-divergence-base base-mixedweb.bin -c 2048 --chunks 60 -ngl 99

Mean KLD and Same top p in that output are the two numbers in the tables above; 99.0% KLD is the worst-1 % column. Every file on this page — ours, Unsloth's, ISTA-DASLab's, AtomicChat's — was measured with those exact commands, same build, same base files, same 60 × 2048 tokens.

Head to head with Unsloth, identical bytes

Same tensor types, same file size, same speed, same memory. What differs is how the weights inside each block were chosen — a different calibration and a different encoder.

held-out neutral web wikitext-2 worst 1% tokens (held-out)
UD-Q4_K_XL 0.0117 0.0087 0.0122 0.088
AP-Q4_K_XL 0.0111 (−4.4%) 0.0083 (−4.5%) 0.0118 0.082 (−7.2%)
UD-Q4_K_M 0.0153 0.0108 0.0139 0.122
AP-Q4_K_M 0.0146 (−4.9%) 0.0104 (−3.7%) 0.0150 0.111 (−8.8%)
UD-IQ4_XS 0.0276 0.0186 0.0252 0.233
AP-IQ4_XS 0.0255 (−7.6%) 0.0181 (−2.4%) 0.0243 (−3.8%) 0.211 (−9.3%)
UD-Q3_K_XL 0.0421 0.0270 0.0337 0.366
AP-Q3_K_XL 0.0339† (−19.5%) 0.0255 (−5.8%) 0.0352 0.280 (−23.5%)
UD-IQ3_S 0.0617 0.0404 0.0470 0.553
AP-IQ3_S 0.0492† (−20.3%) 0.0383 (−5.1%) 0.0506 0.420 (−24.0%)

Held-out gains: 3.2σ, 3.2σ, 5.7σ, 16σ and 18σ (Q4_K_XL / Q4_K_M / IQ4_XS / Q3_K_XL / IQ3_S). Neutral-web gains: 3.7σ at Q4_K_XL, 3.4σ at Q3_K_XL, 3.1σ at IQ3_S; 2.4σ at Q4_K_M. Wikitext-2 is a statistical tie (under 1.5σ) at every tier except IQ3_S, where unsloth is 7% closer (2.1σ).

Top-1 match with BF16 on held-out text: 92.8% · 92.2% · 90.4% · 90.1% · 88.6% · 87.7% · 85.8% (Q4_K_XL / Q4_K_M / IQ4_XS / Q3_K_XL / IQ3_S / IQ3_XS / IQ3_XXS).

The field near these sizes

file size held-out neutral web wikitext-2
AP-Q4_K_XL 16.35 GiB 0.0111 0.0083 0.0118
unsloth UD-Q4_K_XL 16.35 GiB 0.0117 0.0087 0.0122
AP-Q4_K_M 15.33 GiB 0.0146 0.0104 0.0150
unsloth UD-Q4_K_M 15.33 GiB 0.0153 0.0108 0.0139
AtomicChat AD-IQ4_XS-IQ3_S 13.45 GiB 0.0335 0.0234 0.0384
AP-IQ4_XS 13.27 GiB 0.0255 0.0181 0.0243
unsloth UD-IQ4_XS 13.27 GiB 0.0276 0.0186 0.0252
AtomicChat AD-IQ3_S 12.89 GiB 0.0441 0.0303 0.0441
AP-Q3_K_XL 12.24 GiB 0.0339† 0.0255 0.0352
unsloth UD-Q3_K_XL 12.24 GiB 0.0421 0.0270 0.0337
ISTA-DASLab GSQ-RCO-IQ3_S 11.29 GiB 0.0594 0.0432 0.0665
AP-IQ3_S 11.21 GiB 0.0492† 0.0383 0.0506
unsloth UD-IQ3_S 11.21 GiB 0.0617 0.0404 0.0470
AP-IQ3_XS 10.70 GiB 0.0604† 0.0498 0.0667
AtomicChat AD-IQ2_S 10.38 GiB 0.1187 0.0858 0.1062
AP-IQ3_XXS 10.00 GiB 0.0832† 0.0676 0.0869
ISTA-DASLab GSQ-RCO-IQ3_XXS 9.73 GiB 0.1123 0.0824 0.1063
ISTA-DASLab GSQ-RCO-IQ2_S 8.95 GiB 0.1541 0.1143 0.1470
AP-IQ2_S 8.95 GiB 0.1239† 0.0983 0.1204
ISTA-DASLab GSQ-RCO-IQ2_XS 8.17 GiB 0.2263 0.1724 0.2036

ISTA rows are their -mtp builds (draft head included, same as ours).

† AP-IQ2_S (2026-09-28) and AP-IQ3_XXS, AP-IQ3_XS, AP-IQ3_S and AP-Q3_K_XL (2026-09-29) were rebuilt with an extended calibration set that shares 1.4% of its 12-grams with the held-out corpus. Their held-out figures are therefore measured on the 95% of held-out text with zero overlap. That subset is slightly harder: every file we scored on both reads about 3% higher on it than on the full set (ISTA GSQ-RCO-IQ2_S 0.1541 → 0.1584, the previous AP-IQ2_S 0.1418 → 0.1460), so comparing it with the other rows' full held-out figures is conservative. The extended calibration moves precision toward technical and code text: against the previous files, held-out KL drops 15–18% while wikitext-2 reads up to 9% higher.

KL is fidelity to Qwen’s predictions, not a task leaderboard. Downstream evals are next. Until then the claim is narrow and checkable: at every size we ship, you are closer to the original than the same-size alternative.


Running

Use the sampling settings from the Qwen3.8-27B model card. Thinking is on by default. To turn it off per request, send "chat_template_kwargs": {"enable_thinking": false}.

llama-server -hf agentionai/Qwen3.8-27B-AP-GGUF:IQ4_XS \
  --jinja -ngl 999 -fa on -c 32768 -ctk q8_0 -ctv q8_0

Keep the KV cache at q8_0 or f16. A 4-bit value cache makes long reasoning traces degenerate into repetition on this model family.

LM Studio: search for agentionai/Qwen3.8-27B-AP-GGUF and pick a tier. Ollama: ollama run hf.co/agentionai/Qwen3.8-27B-AP-GGUF:IQ4_XS

llama-server -hf agentionai/Qwen3.8-27B-AP-GGUF:IQ4_XS \
  --mmproj mmproj-BF16.gguf --jinja -ngl 999 -fa on -c 32768 -ctk q8_0 -ctv q8_0

Built with

The tiers are built and verified with our own Rust tooling agention-infer. Every tier is measured against BF16 on all three corpora, and shipped only if it beats the same-size alternative. unsloth's dynamic type maps underpin AP-Q4_K_XL, AP-Q4_K_M, AP-IQ4_XS, AP-Q3_K_XL and AP-IQ3_S, and we thank them for that work. AP-IQ3_XS and AP-IQ3_XXS sit at sizes no one else ships and use our own per-tensor allocation, so they are compared against the nearest published files above and below them. AP-IQ2_S is built on ISTA-DASLab's GSQ-RCO type map for that size, re-encoded with our calibration and tooling, and we thank them for that work. imatrix-mixed-v2.gguf, the importance matrix from our calibration pass, is included for anyone building their own quants. Since 2026-09-28 AP-IQ2_S, and since 2026-09-29 AP-IQ3_XXS, AP-IQ3_XS, AP-IQ3_S and AP-Q3_K_XL, use an extended calibration set that adds agentic coding transcripts; its imatrix is not published.

Calibration is ordinary public text: the Bartowski and Thireus imatrix corpora, plus a one-third share of whole articles from the wikitext-2 train split — never the test split, and the 12-gram overlap with each of the three evaluation corpora above is 0, 0 and 6 (section headings and unit conversions). 6,012 documents, 4.67 M characters, md5 b8e3269f085f500ab307b0cc977126b0, 1,200 × 512 tokens through the BF16 model. The calibration is not the advantage here; anyone can build on the same text.

Does it hold up in real use? A one-shot coding test

Same prompt, same five seeds for every file, first answer only, no retries and no fixes. Each page opened in Chrome and judged by eye.

file size working pages looped (no answer)
AP-IQ2_S 8.95 GiB 4 / 5 0
ISTA-DASLab GSQ-RCO-IQ2_S 8.95 GiB 2 / 5 1
ByteShape IQ3_XXS-2.88bpw 9.24 GiB 0 / 5 3

Five trials per file is a small sample; read it as a sanity check that the lower KL carries over, not as a benchmark.

More tiers follow as they clear the same bar.

Support AgentionAI

These quants are released freely. If they save you VRAM or make Qwen more useful, you can buy me a coffee or some GPU time and sponsor continued quantization and benchmarking on GitHub. AgentionAi is a one person team and can use your help.