WhiskyAKM/K2-Horizon-0.9B-NVFP4-GGUF

🤗 Hugging Face sourcetext-generationapache-2.0900M activated635 MBGGUF✓ 1 checksumupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo WhiskyAKM/K2-Horizon-0.9B-NVFP4-GGUF ./model-folder
Needs a seeder →

K2-Horizon-0.9B NVFP4 — GGUF

NVFP4 GGUF conversions derived from the K2-Horizon-0.9B model family. This repository provides NVFP4 (NVIDIA 4-bit floating-point) GGUF files, making the model usable with llama.cpp and other GGUF-compatible inference engines.

Model Overview

K2-Horizon-0.9B is the compact dense member of the K2-Horizon family: a 0.9B-class decoder-only model with a 128K context window. It is a compact reasoning model evaluated across mathematics, coding, science, and tool-use benchmarks.

Model Architecture

Property Value
Architecture K2-Horizon-0.9B
Parameters 0.9B
Context Length 128K tokens (131,072)
Supported Modalities Text
Type Dense

Benchmark Results

Benchmark K2-Horizon-0.9B Qwen3.5-0.8B OpenBMB-1B Qwen3.5-2B
Math
AIME 2025 41.7 1.0 40.4 34.2
AIME 2026 48.5 0.2 40.4 38.8
HMMT Feb 2026 25.8 0.6 23.3 22.7
Scientific Reasoning
GPQA Diamond 27.3 11.9 26.3 54.9
Coding
HumanEval+ 79.9 16.5 65.2 75.6
MBPP+ 68.0 35.4 60.6 67.7
LiveCodeBench v6 37.4 6.6 33.5 29.8
Agents
BFCL v4 28.0 25.3 25.2 43.6

Scores in %. Bold marks the best score in each row. Benchmark results are sourced from the original K2-Horizon-0.9B model description.

GGUF Files

File Format Description
K2-Horizon-0.9B-nvfp4.gguf NVFP4 NVFP4 quantized model

Usage

llama.cpp (CLI)

# Run inference
./llama-cli \
  -m K2-Horizon-0.9B-nvfp4.gguf \
  --temp 1.0 --top-p 0.95

llama-server (OpenAI-compatible API)

# Start the server
./llama-server \
  -m K2-Horizon-0.9B-nvfp4.gguf \
  --port 8080

Key Features

  • Compact reasoning model. A 0.9B-class dense model evaluated across mathematics, coding, science, and tool-use benchmarks.
  • 128K context. Supports up to 131,072 tokens with YaRN RoPE scaling.
  • Multi-teacher distillation. Trained with domain teachers for math and code, STEM, and instruction following.
  • NVFP4 Quantization. Optimized for NVIDIA Blackwell and compatible hardware.

Citation

@misc{k2horizon2026,
  title  = {Introducing K2 Horizon: Frontier Performance, Radically Open},
  author = {{IFM Team}},
  year   = {2026},
  url    = {https://ifm.ai/blog/k2/},
}