TL;DR: VibeThinker-3B at 1.16 GB (3-bit) or 1.35 GB (3.5-bit), compared with 6.18 GB in BF16.
VibeThinker-3B · GSQ-RCO GGUFs
Non-uniform GGUF quantizations produced with GSQ and RCO.
Independent community reproduction. These files were produced by a third party using the published GSQ and RCO methods. They are not an IST-DASLab release and carry no endorsement from the authors of either paper.
Overview
VibeThinker-3B is a dense 3.09B-parameter reasoning model from WeiboAI.
| Method | Description |
|---|---|
| GSQ (Gumbel-Softmax Quantization, paper, code) | Post-training scalar quantization that jointly learns per-coordinate grid assignments and per-group scales through a Gumbel-Softmax relaxation. |
| RCO (Riemannian Constrained Optimization, paper, code) | Assigns one of K quantization types to each of N tensors under an exact total size budget, reformulated as a smooth Riemannian manifold in logit space. |
Both methods were developed at the Deep Algorithms and Systems Lab (DASLab), Institute of Science and Technology Austria.
MMLU-Pro
No reasoning; 2,048-token context; no generated tokens. Each build scores 2,000 questions, stratified across all 14 categories with a fixed seed. Scoring is zero-shot, without a chat template: choose the answer with the highest log likelihood of its full letter continuation ( A, B, etc.).
| Build | Correct | Accuracy |
|---|---|---|
| BF16 reference | 220/2,000 | 11.00% |
| GSQ-RCO 3-bit | 215/2,000 | 10.75% |
| GSQ-RCO 3.5-bit | 216/2,000 | 10.80% |
All three scores are close to chance (11.19% for this subset). The BF16 reference chose A on 1,616 of 2,000 questions. Question IDs and scores.
Results
Perplexity covers 4,088 next-token predictions across eight 1,024-token contexts of held-out text. Lower is better.
| Build | Size | PPL |
|---|---|---|
| BF16 reference | 6.178 GB | 92.5820 |
| GSQ-RCO 3-bit | 1.157 GB | 113.4035 |
| Native imatrix 3-bit | 1.153 GB | 134.2970 |
| Matched initializer 3-bit | 1.157 GB | 184.0001 |
| GSQ-RCO 3.5-bit | 1.350 GB | 77.2388 |
| Native imatrix 3.5-bit | 1.350 GB | 129.1760 |
| Matched initializer 3.5-bit | 1.350 GB | 104.4452 |
Matched initializers have the same tensor types and file sizes as the GSQ builds, with weights from before GSQ training. Native imatrix controls use their own allocations.
GSM8K scores eight questions by exact numerical matching after ####, with a 2,048-token output cap. IFEval scores 16 prompts with the strict checker and a 1,024-token output cap. Both use the source thinking template, temperature 0 and an 8,192-token context.
| Build | GSM8K | GSM8K at token cap | IFEval strict | IFEval at token cap |
|---|---|---|---|---|
| BF16 reference | 7/8 | 0/8 | 2/16 | 15/16 |
| GSQ-RCO 3-bit | 2/8 | 2/8 | 0/16 | 16/16 |
| GSQ-RCO 3.5-bit | 5/8 | 0/8 | 1/16 | 15/16 |
IFEval checks the text after </think>. If the thinking block never closes, it scores an empty answer. One of the BF16 reference's two passes reached the output cap.
Usage
llama.cpp
Tested with unpatched upstream llama.cpp at 58367713.
hf download pfeifferj/VibeThinker-3B-GSQ-RCO-GGUF VibeThinker-3B-GSQ-RCO-3.5bit.gguf --local-dir .
llama-server -m VibeThinker-3B-GSQ-RCO-3.5bit.gguf \
-ngl all -c 8192 --jinja --temp 1.0 --top-p 0.95 --top-k 0
Both GGUFs include the source tokenizer and chat template. The example uses WeiboAI's recommended sampling settings. The authors advise against tool-calling and coding-agent use.
Quantization procedure
Source: WeiboAI/VibeThinker-3B at 77bd2cce. Both files mix Q2_0, Q2_K, Q3_K and Q4_K across the core matrices. Tied embeddings use Q4_K; norms and biases use F32.
- Calibrate all 252 core matrices across 36 layers on one million tokens: 80% OpenThoughts, 15% calibration_mixture and 5% FineWeb. Validation and RCO use separate pools. Each layer retains 4,096 training activation rows and 1,024 validation rows.
- Run 80 GSQ updates per layer and format, sampling 256 activation rows per update. Optimize the quantization codes while keeping native scales and offsets fixed. Select the packed weights with the lowest validation error, including the initializer.
- Run 50 RCO steps per size, minimizing teacher-to-student KL on four calibration sequences. Select the allocation on four separate validation sequences, then assemble the chosen packed tensors into GGUF without requantizing them.
Quantization settings and tensor assignments.
Citation
Please cite the release, base model and methods:
This release · DOI: 10.57967/hf/10577
@misc{vibethinkergsqrco2026,
title = {VibeThinker-3B GSQ-RCO GGUF quantizations},
author = {Josephine Pfeiffer},
year = {2026},
publisher = {Hugging Face},
doi = {10.57967/hf/10577},
howpublished = {\url{https://huggingface.co/pfeifferj/VibeThinker-3B-GSQ-RCO-GGUF}}
}
Base model
@misc{xu2026vibethinker3bexploringfrontierverifiable,
title = {VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models},
author = {Sen Xu and Shixi Liu and Wei Wang and Jixin Min and Yingwei Dai and Zhibin Yin and Yirong Chen and Xin Zhou and Junlin Zhang},
year = {2026},
eprint = {2606.16140},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2606.16140}
}
Methods
@article{gsq2026,
title = {GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling},
author = {Dadgarnia, Alireza and Tabesh, Soroush and Nikdan, Mahdi and Helcig, Michael and Kurtic, Eldar and Kleinegger, Maximilian and Alistarh, Dan},
journal= {arXiv preprint arXiv:2604.18556},
year = {2026}
}
@article{rco2026,
title = {Model Compression with Exact Budget Constraints via Riemannian Manifolds},
author = {Helcig, Michael and Alistarh, Dan},
journal= {arXiv preprint arXiv:2605.00649},
year = {2026}
}
Acknowledgements
Thanks to WeiboAI for VibeThinker-3B and to the Deep Algorithms and Systems Lab (DASLab) for GSQ, RCO and their reference implementations.
License
These weights use the base model's MIT license. LICENSE is copied from WeiboAI's VibeThinker repository. The GSQ and RCO code has its own licenses.