vcruz305/Ternary-Bonsai-27B-GGUF

🤗 Hugging Face sourcetext-generationapache-2.027B activated116 GBGGUF✓ 9 checksumsupdated today
Needs seeder →

Ternary-Bonsai-27B — Standard GGUF Ladder (Q3_K_M → Q8_0)

Imatrix GGUF quants of prism-ml's Ternary-Bonsai-27B (Qwen3.6-27B backbone, 262K context, hybrid attention, vision), covering the quality tiers between the official QAT release and F16.

Which repo should you use?

If you have ≤8GB: use prism-ml's official QAT quants, not these.

  • Q2_0 / PQ2_0 (7.2GB), Q2_g64 (7.6GB) — 95% of F16 quality, quantization-aware-trained with custom kernels. At those sizes they beat anything post-training quantization can produce, including anything in this repo.

This repo covers the gap above them: the standard ladder for 12–32GB setups, quantized from prism-ml's own F16 GGUF with an importance matrix (63KB coding/reasoning calibration corpus).

Quants

File Size Measured (GB10, 273GB/s)
Q8_0 28.6 GB 7.8 tok/s
Q6_K 22.1 GB 9.1 tok/s
Q5_K_M 19.2 GB 10.4 tok/s
Q4_K_M 16.6 GB 12.3 tok/s
IQ4_XS 15.1 GB 14.0 tok/s
Q3_K_M 13.3 GB 13.6 tok/s

All tiers individually smoke-tested (coherent code generation, chat template engages via --jinja). Q8_0 quantized without imatrix (unneeded at 8-bit); all others use the included Ternary-Bonsai-27B.imatrix. Decode rates scale with memory bandwidth — a 936GB/s GPU (RTX 3090-class) will run ~3× these numbers.

Both official mmproj files are included for vision (--mmproj Ternary-Bonsai-27B-mmproj-Q8_0.gguf).

DSpark drafter: tested, NOT recommended with these quants

We verified prism-ml's Ternary DSpark drafter against this repo's Q4_K_M using their llama.cpp fork (branch prism):

Config tok/s Draft acceptance
Q4_K_M base 11.9 —
Q4_K_M + DSpark drafter 11.9–12.7 54% (217/404)

Acceptance drops to ~54% (vs 79% for the Bonsai-27B pairing, where it delivers +31%) — at that rate the draft overhead roughly cancels the win on bandwidth-bound hardware. The drafter appears to track the ternary QAT weights more tightly than the F16 latent this ladder is quantized from. On high-bandwidth GPUs (where drafting is cheaper) it may still net positive — measure before adopting, and check timings.draft_n_accepted in any API response to verify. For guaranteed drafter gains at small sizes, use prism-ml's official QAT Q2_0 + drafter combination.

Provenance

  • Source: prism-ml/Ternary-Bonsai-27B-gguf F16 (53.8GB) — quantized directly from their official F16 GGUF, no re-conversion
  • llama.cpp: mainline, commit cecbf5fb0 (quantization + baseline numbers); PrismML fork 62061f91 branch prism (drafter test)
  • imatrix: 63KB ChatML coding/debugging/reasoning corpus, 4096 ctx
  • LICENSE/NOTICE carried over from the upstream repo (Apache-2.0)
  • Quantized on NVIDIA DGX Spark (GB10, aarch64) — day-zero release, same-day as the upstream drop
  • Companion repo: vcruz305/Bonsai-27B-GGUF — the non-ternary ladder, where the DSpark drafter pairing IS verified worthwhile (+31%)

Report issues in the community tab — smoke-test failures, incoherence, or numbers that don't reproduce. Community benchmark reports welcome (include hardware, backend, and full launch command).