WhiskyAKM/K2-Horizon-MoVA-36B-A4B-NVFP4-GGUF

🤗 Hugging Face 来源text-generationapache-2.0激活 4B21 GBGGUF✓ 1 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo WhiskyAKM/K2-Horizon-MoVA-36B-A4B-NVFP4-GGUF ./model-folder
需要做种者 →

K2-Horizon-MoVA-36B-A4B NVFP4 — GGUF

NVFP4 GGUF conversions derived from the K2-Horizon-MoVA-36B-A4B model family. This repository provides NVFP4 (NVIDIA 4-bit floating-point) GGUF files, making the model usable with llama.cpp and other GGUF-compatible inference engines.

Model Overview

K2-Horizon-MoVA-36B-A4B is the sparse member of the K2-Horizon family: a Mixture-of-Experts model with Mixture-of-Values attention (MoVA) that stores 36B parameters and runs 4B per token. It delivers frontier-class results with high efficiency in agentic and reasoning tasks.

Model Architecture

Property Value
Architecture MoE with Mixture-of-Values (MoVA)
Total Parameters 36B
Activated Parameters 4B
Context Length 512K tokens (524,288)
Supported Modalities Text
Type Sparse (MoE)

Benchmark Results

Benchmark K2-Horizon-MoVA-36B-A4B Nemotron 3 Ultra Nemotron 3 Super G9v3-39A5B Qwen3.6-35B-A3B Muse Glimmer-30B Gemma 4 31B-it
# Params 36B 550B 120B 39B 35B 30B 31B
# Active Params 4B 55B 12B 5B 3B 30B 31B
Agents
tau3-Banking 26.8 14.2 10.3 22.1 9.3 23.5 14.8
Coding
Terminal-Bench 2.1 58.6 53.9 38.6 32.6 44.9 51.7 43.4
SciCode 38.9 39.9 36.0 34.0 35.8 43.6 43.4
Scientific Reasoning
Humanity's Last Exam 25.2 28.4 20.8 17.5 22.2 22.0 23.6
GPQA Diamond 80.8 86.7 80.0 80.5 84.1 83.5 85.7
CritPt 2.1 3.1 3.1 0.3 0.3 2.6 1.4
General
AA-LCR 66.3 71.0 60.3 62.0 66.7 80.0 68.3
AA-Omniscience Acc. 18.8 22.6 24.3 14.9 18.8 27.0 20.0
AA-Omniscience Non-Hall. 69.2 70.3 13.0 87.0 49.5 18.1 15.0

Scores in %. Bold marks the best score in each row. Results are sourced from the original K2-Horizon-MoVA-36B-A4B model description.*

GGUF Files

File Format Description
k2-horizon-mova-36b-a4b-nvfp4.gguf NVFP4 NVFP4 quantized model

Usage

llama.cpp (CLI)

# Run inference
./llama-cli \
  -m k2-horizon-mova-36b-a4b-nvfp4.gguf \
  --temp 1.0 --top-p 0.95

llama-server (OpenAI-compatible API)

# Start the server
./llama-server \
  -m k2-horizon-mova-36b-a4b-nvfp4.gguf \
  --port 8080

Key Features

  • High Efficiency MoE. Achieves frontier-class results with only 4B active parameters, outscoring many dense models significantly larger.
  • 512K Context. Native 524,288-token context window supported from midtraining onward.
  • NVFP4 Quantization. Optimized for NVIDIA Blackwell and compatible hardware for high-performance inference.
  • Reasoning Capabilities. Strong performance in scientific reasoning, coding, and agentic tool use.

Citation

@misc{k2horizon2026,
  title  = {Introducing K2 Horizon: Frontier Performance, Radically Open},
  author = {{IFM Team}},
  year   = {2026},
  url    = {https://ifm.ai/blog/k2/},
}