Ornith-1.5-9B GSQ-RCO GGUF
Mixed-precision GSQ-RCO GGUF quantizations of ornith-ai/Ornith-1.5-9B, optimized for efficient, high-fidelity local inference in llama.cpp.
Benchmarks
Why GSQ-RCO?
Standard quantization severely degrades hybrid SSM-Transformer models (e.g. stock IQ3_M perplexity drops to ~75.2).
GSQ-RCO solves this through two targeted techniques:
- RCO: Keeps all 48 critical SSM parameters in unquantized BF16.
- GSQ: Optimizes per-layer precision guided by Hessian sensitivity calibration.
GGUF Files
| File | bpw | Size |
|---|---|---|
| Ornith-1.5-9B-GSQ-RCO-IQ2_XS.gguf | 2.50 | 2.80 GB |
| Ornith-1.5-9B-GSQ-RCO-IQ2_S.gguf | 2.75 | 3.08 GB |
| Ornith-1.5-9B-GSQ-RCO-IQ3_XXS.gguf | 3.00 | 3.36 GB |
| Ornith-1.5-9B-GSQ-RCO-IQ3_XS.gguf | 3.25 | 3.64 GB |
| Ornith-1.5-9B-GSQ-RCO-IQ3_S.gguf | 3.50 | 3.92 GB |
| Ornith-1.5-9B-GSQ-RCO-IQ3_M.gguf (Recommended) | 3.75 | 4.20 GB |
Each quantization also ships an optional -mtp build that carries the model's Multi-Token Prediction head (layer 32) for speculative decoding in supported runtimes.
Attribution & License
- Distributed under the MIT License, matching upstream ornith-ai/Ornith-1.5-9B.
- Reference GGUF release: ornith-ai/Ornith-1.5-9B-GGUF.