abenzerps/Ornith-1.5-9B-GSQ-RCO-GGUF

Verified creator abenzerps verified
🤗 Hugging Face sourcetext-generationmit9B activated44 GBGGUF✓ 13 checksumsupdated today
Needs seeder →

Ornith-1.5-9B GSQ-RCO GGUF

Mixed-precision GSQ-RCO GGUF quantizations of ornith-ai/Ornith-1.5-9B, optimized for efficient, high-fidelity local inference in llama.cpp.

Benchmarks

Why GSQ-RCO?

Standard quantization severely degrades hybrid SSM-Transformer models (e.g. stock IQ3_M perplexity drops to ~75.2).

GSQ-RCO solves this through two targeted techniques:

  • RCO: Keeps all 48 critical SSM parameters in unquantized BF16.
  • GSQ: Optimizes per-layer precision guided by Hessian sensitivity calibration.

GGUF Files

Each quantization also ships an optional -mtp build that carries the model's Multi-Token Prediction head (layer 32) for speculative decoding in supported runtimes.

Attribution & License