Ternary-Bonsai-27B — Standard GGUF Ladder (Q3_K_M → Q8_0)
Imatrix GGUF quants of prism-ml's Ternary-Bonsai-27B (Qwen3.6-27B backbone, 262K context, hybrid attention, vision), covering the quality tiers between the official QAT release and F16.
Which repo should you use?
If you have ≤8GB: use prism-ml's official QAT quants, not these.
- Q2_0 / PQ2_0 (7.2GB), Q2_g64 (7.6GB) — 95% of F16 quality, quantization-aware-trained with custom kernels. At those sizes they beat anything post-training quantization can produce, including anything in this repo.
This repo covers the gap above them: the standard ladder for 12–32GB setups, quantized from prism-ml's own F16 GGUF with an importance matrix (63KB coding/reasoning calibration corpus).
Quants
| File | Size | Measured (GB10, 273GB/s) |
|---|---|---|
| Q8_0 | 28.6 GB | 7.8 tok/s |
| Q6_K | 22.1 GB | 9.1 tok/s |
| Q5_K_M | 19.2 GB | 10.4 tok/s |
| Q4_K_M | 16.6 GB | 12.3 tok/s |
| IQ4_XS | 15.1 GB | 14.0 tok/s |
| Q3_K_M | 13.3 GB | 13.6 tok/s |
All tiers individually smoke-tested (coherent code generation, chat template engages via --jinja). Q8_0 quantized without imatrix (unneeded at 8-bit); all others use the included Ternary-Bonsai-27B.imatrix. Decode rates scale with memory bandwidth — a 936GB/s GPU (RTX 3090-class) will run ~3× these numbers.
Both official mmproj files are included for vision (--mmproj Ternary-Bonsai-27B-mmproj-Q8_0.gguf).
DSpark drafter: tested, NOT recommended with these quants
We verified prism-ml's Ternary DSpark drafter against this repo's Q4_K_M using their llama.cpp fork (branch prism):
| Config | tok/s | Draft acceptance |
|---|---|---|
| Q4_K_M base | 11.9 | — |
| Q4_K_M + DSpark drafter | 11.9–12.7 | 54% (217/404) |
Acceptance drops to ~54% (vs 79% for the Bonsai-27B pairing, where it delivers +31%) — at that rate the draft overhead roughly cancels the win on bandwidth-bound hardware. The drafter appears to track the ternary QAT weights more tightly than the F16 latent this ladder is quantized from. On high-bandwidth GPUs (where drafting is cheaper) it may still net positive — measure before adopting, and check timings.draft_n_accepted in any API response to verify. For guaranteed drafter gains at small sizes, use prism-ml's official QAT Q2_0 + drafter combination.
Provenance
- Source:
prism-ml/Ternary-Bonsai-27B-ggufF16 (53.8GB) — quantized directly from their official F16 GGUF, no re-conversion - llama.cpp: mainline, commit
cecbf5fb0(quantization + baseline numbers); PrismML fork62061f91branchprism(drafter test) - imatrix: 63KB ChatML coding/debugging/reasoning corpus, 4096 ctx
- LICENSE/NOTICE carried over from the upstream repo (Apache-2.0)
- Quantized on NVIDIA DGX Spark (GB10, aarch64) — day-zero release, same-day as the upstream drop
- Companion repo: vcruz305/Bonsai-27B-GGUF — the non-ternary ladder, where the DSpark drafter pairing IS verified worthwhile (+31%)
Report issues in the community tab — smoke-test failures, incoherence, or numbers that don't reproduce. Community benchmark reports welcome (include hardware, backend, and full launch command).