K2-Horizon-MoVA-36B-A4B NVFP4 — GGUF
NVFP4 GGUF conversions derived from the K2-Horizon-MoVA-36B-A4B model family. This repository provides NVFP4 (NVIDIA 4-bit floating-point) GGUF files, making the model usable with llama.cpp and other GGUF-compatible inference engines.
Model Overview
K2-Horizon-MoVA-36B-A4B is the sparse member of the K2-Horizon family: a Mixture-of-Experts model with Mixture-of-Values attention (MoVA) that stores 36B parameters and runs 4B per token. It delivers frontier-class results with high efficiency in agentic and reasoning tasks.
Model Architecture
| Property | Value |
|---|---|
| Architecture | MoE with Mixture-of-Values (MoVA) |
| Total Parameters | 36B |
| Activated Parameters | 4B |
| Context Length | 512K tokens (524,288) |
| Supported Modalities | Text |
| Type | Sparse (MoE) |
Benchmark Results
| Benchmark | K2-Horizon-MoVA-36B-A4B | Nemotron 3 Ultra | Nemotron 3 Super | G9v3-39A5B | Qwen3.6-35B-A3B | Muse Glimmer-30B | Gemma 4 31B-it |
|---|---|---|---|---|---|---|---|
| # Params | 36B | 550B | 120B | 39B | 35B | 30B | 31B |
| # Active Params | 4B | 55B | 12B | 5B | 3B | 30B | 31B |
| Agents | |||||||
| tau3-Banking | 26.8 | 14.2 | 10.3 | 22.1 | 9.3 | 23.5 | 14.8 |
| Coding | |||||||
| Terminal-Bench 2.1 | 58.6 | 53.9 | 38.6 | 32.6 | 44.9 | 51.7 | 43.4 |
| SciCode | 38.9 | 39.9 | 36.0 | 34.0 | 35.8 | 43.6 | 43.4 |
| Scientific Reasoning | |||||||
| Humanity's Last Exam | 25.2 | 28.4 | 20.8 | 17.5 | 22.2 | 22.0 | 23.6 |
| GPQA Diamond | 80.8 | 86.7 | 80.0 | 80.5 | 84.1 | 83.5 | 85.7 |
| CritPt | 2.1 | 3.1 | 3.1 | 0.3 | 0.3 | 2.6 | 1.4 |
| General | |||||||
| AA-LCR | 66.3 | 71.0 | 60.3 | 62.0 | 66.7 | 80.0 | 68.3 |
| AA-Omniscience Acc. | 18.8 | 22.6 | 24.3 | 14.9 | 18.8 | 27.0 | 20.0 |
| AA-Omniscience Non-Hall. | 69.2 | 70.3 | 13.0 | 87.0 | 49.5 | 18.1 | 15.0 |
Scores in %. Bold marks the best score in each row. Results are sourced from the original K2-Horizon-MoVA-36B-A4B model description.*
GGUF Files
| File | Format | Description |
|---|---|---|
k2-horizon-mova-36b-a4b-nvfp4.gguf |
NVFP4 | NVFP4 quantized model |
Usage
llama.cpp (CLI)
# Run inference
./llama-cli \
-m k2-horizon-mova-36b-a4b-nvfp4.gguf \
--temp 1.0 --top-p 0.95
llama-server (OpenAI-compatible API)
# Start the server
./llama-server \
-m k2-horizon-mova-36b-a4b-nvfp4.gguf \
--port 8080
Key Features
- High Efficiency MoE. Achieves frontier-class results with only 4B active parameters, outscoring many dense models significantly larger.
- 512K Context. Native 524,288-token context window supported from midtraining onward.
- NVFP4 Quantization. Optimized for NVIDIA Blackwell and compatible hardware for high-performance inference.
- Reasoning Capabilities. Strong performance in scientific reasoning, coding, and agentic tool use.
Citation
@misc{k2horizon2026,
title = {Introducing K2 Horizon: Frontier Performance, Radically Open},
author = {{IFM Team}},
year = {2026},
url = {https://ifm.ai/blog/k2/},
}