WhiskyAKM/K2-Horizon-MoVA-36B-A4B-NVFP4-GGUF

🤗 Hugging Face sourcetext-generationapache-2.04B activated21 GBGGUF✓ 1 checksumupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo WhiskyAKM/K2-Horizon-MoVA-36B-A4B-NVFP4-GGUF ./model-folder
Needs a seeder →

K2-Horizon-MoVA-36B-A4B NVFP4 — GGUF

NVFP4 GGUF conversions derived from the K2-Horizon-MoVA-36B-A4B model family. This repository provides NVFP4 (NVIDIA 4-bit floating-point) GGUF files, making the model usable with llama.cpp and other GGUF-compatible inference engines.

Model Overview

K2-Horizon-MoVA-36B-A4B is the sparse member of the K2-Horizon family: a Mixture-of-Experts model with Mixture-of-Values attention (MoVA) that stores 36B parameters and runs 4B per token. It delivers frontier-class results with high efficiency in agentic and reasoning tasks.

Model Architecture

Property Value
Architecture MoE with Mixture-of-Values (MoVA)
Total Parameters 36B
Activated Parameters 4B
Context Length 512K tokens (524,288)
Supported Modalities Text
Type Sparse (MoE)

Benchmark Results

Benchmark K2-Horizon-MoVA-36B-A4B Nemotron 3 Ultra Nemotron 3 Super G9v3-39A5B Qwen3.6-35B-A3B Muse Glimmer-30B Gemma 4 31B-it
# Params 36B 550B 120B 39B 35B 30B 31B
# Active Params 4B 55B 12B 5B 3B 30B 31B
Agents
tau3-Banking 26.8 14.2 10.3 22.1 9.3 23.5 14.8
Coding
Terminal-Bench 2.1 58.6 53.9 38.6 32.6 44.9 51.7 43.4
SciCode 38.9 39.9 36.0 34.0 35.8 43.6 43.4
Scientific Reasoning
Humanity's Last Exam 25.2 28.4 20.8 17.5 22.2 22.0 23.6
GPQA Diamond 80.8 86.7 80.0 80.5 84.1 83.5 85.7
CritPt 2.1 3.1 3.1 0.3 0.3 2.6 1.4
General
AA-LCR 66.3 71.0 60.3 62.0 66.7 80.0 68.3
AA-Omniscience Acc. 18.8 22.6 24.3 14.9 18.8 27.0 20.0
AA-Omniscience Non-Hall. 69.2 70.3 13.0 87.0 49.5 18.1 15.0

Scores in %. Bold marks the best score in each row. Results are sourced from the original K2-Horizon-MoVA-36B-A4B model description.*

GGUF Files

File Format Description
k2-horizon-mova-36b-a4b-nvfp4.gguf NVFP4 NVFP4 quantized model

Usage

llama.cpp (CLI)

# Run inference
./llama-cli \
  -m k2-horizon-mova-36b-a4b-nvfp4.gguf \
  --temp 1.0 --top-p 0.95

llama-server (OpenAI-compatible API)

# Start the server
./llama-server \
  -m k2-horizon-mova-36b-a4b-nvfp4.gguf \
  --port 8080

Key Features

  • High Efficiency MoE. Achieves frontier-class results with only 4B active parameters, outscoring many dense models significantly larger.
  • 512K Context. Native 524,288-token context window supported from midtraining onward.
  • NVFP4 Quantization. Optimized for NVIDIA Blackwell and compatible hardware for high-performance inference.
  • Reasoning Capabilities. Strong performance in scientific reasoning, coding, and agentic tool use.

Citation

@misc{k2horizon2026,
  title  = {Introducing K2 Horizon: Frontier Performance, Radically Open},
  author = {{IFM Team}},
  year   = {2026},
  url    = {https://ifm.ai/blog/k2/},
}