WhiskyAKM/K2-Horizon-3.7B-NVFP4-GGUF

🤗 Hugging Face sourcetext-generationapache-2.03.7B activated3.0 GBGGUF✓ 1 checksumupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo WhiskyAKM/K2-Horizon-3.7B-NVFP4-GGUF ./model-folder
Needs a seeder →

K2-Horizon-3.7B NVFP4 — GGUF

NVFP4 GGUF conversions derived from the K2-Horizon-3.7B model family. This repository provides NVFP4 (NVIDIA 4-bit floating-point) GGUF files, making the model usable with llama.cpp and other GGUF-compatible inference engines.

Model Overview

K2-Horizon-3.7B is the small dense member of the K2-Horizon family: a 3.7B-core decoder-only model with a 512K context window.

Model Architecture

Property Value
Architecture K2-Horizon-3.7B
Parameters 3.7B
Context Length 512K tokens (524,288)
Supported Modalities Text
Type Dense

Benchmark Results

Benchmark K2-Horizon-3.7B Qwen3.5-4B G9v3-3B Granite 4.2-3B Nemotron 3 Nano-4B
Math
HMMT Feb 2026 70.5 61.6 34.1 57.2 34.7
Coding
SWE-bench Verified 68.6 41.2 16.4 32.2 1.8
Scientific Reasoning
GPQA Diamond 65.4 77.1 43.8 55.9 51.3
HLE 12.9 9.9 4.5 6.6 4.9
SciCode 25.9 16.1 17.7 24.9 16.4
Agents
Terminal-Bench 2.1 25.1 25.8 6.0 13.9 3.7
tau3-Banking 17.7 6.8 — 5.6 —
BFCL v4 50.9 55.7 47.9 50.8 36.8

Scores in %. Bold marks the best score in each row. Benchmark results are sourced from the original K2-Horizon-3.7B model description.

GGUF Files

File Format Description
K2-Horizon-3.7B-nvfp4.gguf NVFP4 NVFP4 quantized model

Usage

llama.cpp (CLI)

# Run inference
./llama-cli \
  -m K2-Horizon-3.7B-nvfp4.gguf \
  --temp 1.0 --top-p 0.95

llama-server (OpenAI-compatible API)

# Start the server
./llama-server \
  -m K2-Horizon-3.7B-nvfp4.gguf \
  --port 8080

Key Features

  • Strong small-model baseline. A dense model evaluated on the same agentic, coding, and reasoning benchmarks as the rest of the family.
  • 512K context. Native 524,288-token context from the midtraining stages onward.
  • NVFP4 Quantization. Optimized for NVIDIA Blackwell and compatible hardware.

Citation

@misc{k2horizon2026,
  title  = {Introducing K2 Horizon: Frontier Performance, Radically Open},
  author = {{IFM Team}},
  year   = {2026},
  url    = {https://ifm.ai/blog/k2/},
}