K2-Horizon-7B NVFP4 — GGUF
NVFP4 GGUF conversions derived from the K2-Horizon-7B model family. This repository provides NVFP4 (NVIDIA 4-bit floating-point) GGUF files, making the model usable with llama.cpp and other GGUF-compatible inference engines.
Model Overview
K2-Horizon-7B is the medium dense member of the K2-Horizon family: a 7B-core decoder-only model with a 512K context window. It is a strong dense baseline evaluated across agentic, coding, long-context, and reasoning benchmarks.
Model Architecture
| Property | Value |
|---|---|
| Architecture | K2-Horizon-7B |
| Parameters | 7B |
| Context Length | 512K tokens (524,288) |
| Supported Modalities | Text |
| Type | Dense |
Benchmark Results
| Benchmark | K2-Horizon-7B | Reference 1 | Reference 2 | Reference 3 |
|---|---|---|---|---|
| Math | ||||
| HMMT Feb 2026 | 73.3 | 63.1 (Gemma 4-12B) | 65.7 (Qwen3.5-9B) | 66.5 (Granite 4.2-8B) |
| Coding | ||||
| SWE-bench Verified | 70.6 | 30.6 (Gemma 4-12B) | 47.7 (Granite 4.2-8B) | 50.8 (Qwen3.5-9B) |
| Scientific Reasoning | ||||
| HLE | 18.6 | 9.7 (Granite 4.2-8B) | 14.9 (Qwen3.5-9B) | 15.7 (Gemma 4-12B) |
| SciCode | 31.6 | 27.5 (Qwen3.5-9B) | 28.0 (Mistral Small 4) | 30.4 (Granite 4.2-8B) |
| General | ||||
| LCR | 68.0 | 43.3 (Granite 4.2-8B) | 61.7 (Gemma 4-12B) | 65.3 (Qwen3.5-9B) |
| Coding | ||||
| Terminal-Bench 2.1 | 39.1 | 18.4 (Granite 4.2-8B) | 27.3 (Gemma 4-12B) | 29.2 (Qwen3.5-9B) |
| Agents | ||||
| tau3-Banking | 25.8 | 7.0 (Qwen3.5-9B) | 7.6 (Granite 4.2-8B) | 24.0 (Muse Glimmer-30B) |
| BrowseComp | 59.0 | 53.5 (DeepSeek V4 Flash) | 54.9 (GPT-5) | 56.6 (LongCat Flash) |
Scores in %. Bold marks the best score in each row. Benchmark results are sourced from the original K2-Horizon-7B model description.*
GGUF Files
| File | Format | Description |
|---|---|---|
K2-Horizon-7B-nvfp4.gguf |
NVFP4 | NVFP4 quantized model |
Usage
llama.cpp (CLI)
# Run inference
./llama-cli \
-m K2-Horizon-7B-nvfp4.gguf \
--temp 1.0 --top-p 0.95
llama-server (OpenAI-compatible API)
# Start the server
./llama-server \
-m K2-Horizon-7B-nvfp4.gguf \
--port 8080
Key Features
- Strong dense baseline. A 7B-class dense model evaluated across agentic, coding, long-context, and reasoning benchmarks.
- 512K context. Native 524,288-token context from the midtraining stages onward.
- NVFP4 Quantization. Optimized for NVIDIA Blackwell and compatible hardware.
Citation
@misc{k2horizon2026,
title = {Introducing K2 Horizon: Frontier Performance, Radically Open},
author = {{IFM Team}},
year = {2026},
url = {https://ifm.ai/blog/k2/},
}