pfeifferj/Ornith-1.5-35B-A3B-GSQ-RCO-GGUF

🤗 Hugging Face 来源text-generationmit激活 3B28 GBGGUF✓ 3 个校验和今天更新
需要做种者 →

TL;DR: Text-only Ornith-1.5-35B-A3B at 13.0 GB (3-bit) or 15.2 GB (3.5-bit), down from 69.4 GB in BF16.

Ornith-1.5-35B-A3B · GSQ-RCO GGUFs

3-bit and 3.5-bit GGUFs made with GSQ and RCO.

Independent community reproduction. These files were produced by a third party using the published GSQ and RCO methods. They are not an IST-DASLab release and carry no endorsement from the authors of either paper.


Overview

Ornith-1.5-35B-A3B is a Qwen3.5-family mixture-of-experts model with about 35B parameters and 3B active per token. Both GGUFs include the text model, tokenizer and chat template.

Method Description
GSQ (Gumbel-Softmax Quantization, paper, code) Post-training scalar quantization that jointly learns per-coordinate grid assignments and per-group scales through a Gumbel-Softmax relaxation.
RCO (Riemannian Constrained Optimization, paper, code) Assigns one of K quantization types to each of N tensors under an exact total size budget, reformulated as a smooth Riemannian manifold in logit space.

Both methods were developed at the Deep Algorithms and Systems Lab (DASLab), Institute of Science and Technology Austria.


Results

MMLU-Pro: no reasoning, 2,048-token context, zero generated tokens. The test uses 2,000 questions across all 14 categories. Zero-shot scoring chooses the answer with the highest log likelihood for its full letter continuation ( A, B, etc.), without a chat template. Question IDs and scores.

Perplexity covers 4,088 next-token predictions across eight 1,024-token contexts of held-out text. Lower is better.

Build Size PPL MMLU-Pro
BF16 reference 69.377 GB 3.4112 47.60%
GSQ-RCO 3-bit 12.992 GB 3.5369 43.70%
Native imatrix 3-bit 12.984 GB 3.5320 41.90%
Matched initializer 3-bit 12.992 GB 3.5353 45.35%
GSQ-RCO 3.5-bit 15.163 GB 3.4575 43.20%
Native imatrix 3.5-bit 15.156 GB 3.4795 43.95%
Matched initializer 3.5-bit 15.163 GB 3.4575 44.65%

Perplexity rises 3.7% at 3-bit and 1.4% at 3.5-bit over BF16. Both GSQ builds score below their matched initializers on MMLU-Pro.

Matched initializers use the same tensor formats and file sizes as the GSQ builds, with native quantized weights before GSQ optimization. The native imatrix controls have separate allocations.

GSM8K uses eight questions, thinking enabled and a 2,048-token output cap; scoring checks the numerical answer after ####. IFEval uses 16 prompts, thinking disabled, strict scoring and a 1,024-token cap. Both use temperature 0 and an 8,192-token context.

Build GSM8K IFEval strict IFEval at token cap
BF16 reference 4/8 14/16 2/16
GSQ-RCO 3-bit 8/8 11/16 3/16
GSQ-RCO 3.5-bit 7/8 14/16 1/16

Two BF16 IFEval passes and one 3-bit pass reached the cap. Excluding capped passes gives 12/16, 10/16 and 14/16, respectively. No GSM8K answer reached the cap.


Usage

llama.cpp

Tested with upstream llama.cpp at 58367713.

hf download pfeifferj/Ornith-1.5-35B-A3B-GSQ-RCO-GGUF Ornith-1.5-35B-A3B-GSQ-RCO-3.5bit.gguf --local-dir .

llama-server -m Ornith-1.5-35B-A3B-GSQ-RCO-3.5bit.gguf \
  -ngl 999 -c 8192 -fa on -ctk f32 -ctv f32 --jinja \
  --temp 0.6 --top-p 0.95 --top-k 20 \
  --chat-template-kwargs '{"preserve_thinking":true}'

These are Ornith's recommended sampling settings. Thinking is on by default; disable it with --chat-template-kwargs '{"enable_thinking":false,"preserve_thinking":true}'.


Quantization procedure

Source: ornith-ai/Ornith-1.5-35B-A3B at 10fbf86f. Routed experts mix Q2_K, Q3_K and Q4_K. Embeddings and the output matrix use Q4_K; attention and shared-expert matrices use Q8_0. Routers, norms and remaining convolution and state tensors use F32.

  1. Capture expert activations across 40 layers from one million tokens: 45% calibration_mixture, 40% OpenThoughts and 15% FineWeb. Keep up to 128 training and 64 validation rows per expert.
  2. Run 40 GSQ updates per projection and format with 32 rows per update. Fit gate, up and down outputs separately, using the original gate/up activations for the down projection. Scales and offsets stay fixed; keep the initializer when it validates better.
  3. Run 50 RCO steps per size on four calibration sequences and choose the allocation on four separate validation sequences. Assemble the selected packed tensors into GGUF without requantization. Quantization settings.

Citation

This release · DOI: 10.57967/hf/10579

@misc{ornithgsqrco2026,
  title        = {Ornith-1.5-35B-A3B GSQ-RCO GGUF quantizations},
  author       = {Josephine Pfeiffer},
  year         = {2026},
  publisher    = {Hugging Face},
  doi          = {10.57967/hf/10579},
  howpublished = {\url{https://huggingface.co/pfeifferj/Ornith-1.5-35B-A3B-GSQ-RCO-GGUF}}
}

Base model

@misc{ornith_1_5,
  title  = {{Ornith-1.5}: From Self-Scaffolding to Self-Improvement},
  url    = {https://ornith.ai/ornith_1_5.html},
  author = {{Ornith Team}},
  year   = {2026}
}

Methods

@article{gsq2026,
  title  = {GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling},
  author = {Dadgarnia, Alireza and Tabesh, Soroush and Nikdan, Mahdi and Helcig, Michael and Kurtic, Eldar and Kleinegger, Maximilian and Alistarh, Dan},
  journal= {arXiv preprint arXiv:2604.18556},
  year   = {2026}
}
@article{rco2026,
  title  = {Model Compression with Exact Budget Constraints via Riemannian Manifolds},
  author = {Helcig, Michael and Alistarh, Dan},
  journal= {arXiv preprint arXiv:2605.00649},
  year   = {2026}
}

Acknowledgements

Thanks to the Ornith team for Ornith-1.5 and to the Deep Algorithms and Systems Lab (DASLab) for GSQ, RCO and their reference implementations.


License

The weights use the base model's MIT license. LICENSE comes from the official Ornith repository.