Abiray/K2-Horizon-MoVA-36B-A4B-GGUF

🤗 Hugging Face sourcetext-generationapache-2.04B activated175 GBGGUF✓ 7 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Abiray/K2-Horizon-MoVA-36B-A4B-GGUF ./model-folder
Needs a seeder →

K2-Horizon-MoVA-36B-A4B-GGUF

This repository contains GGUF quantizations of IFM/K2-Horizon-MoVA-36B-A4B, a sparse Mixture-of-Experts model utilizing Mixture-of-Values Attention (MoVA). It features 36B total parameters with only ~4B activated per token, enabling near-small-model inference speeds with 36B-scale intelligence.

[!IMPORTANT] Architecture Support Required: As of early 2026, k2-horizon is a newly introduced architecture not yet merged into upstream mainline llama.cpp. To run or quantize these GGUFs, you must build from the official MBZUAI-IFM fork (model/K2Horizon branch). Running on vanilla llama.cpp will result in unknown model architecture: 'k2-horizon'.


Quantization Overview & Hardware Recommendations

File Name Quant Size Recommended VRAM (Full Offload) Minimum System RAM (CPU) Description
K2-Horizon-MoVA-36B-A4B-Q3_K_S.gguf Q3_K_S 16.4 GB 20 GB (RTX 3090/4090) 24 GB Smallest footprint; lowest memory usage.
K2-Horizon-MoVA-36B-A4B-Q3_K_M.gguf Q3_K_M 17.7 GB 22 GB (RTX 3090/4090) 28 GB Balanced 3-bit quantization.
K2-Horizon-MoVA-36B-A4B-Q4_K_S.gguf Q4_K_S 21.4 GB 24 GB (RTX 3090/4090) 32 GB Great performance/size ratio for 24 GB GPUs.
K2-Horizon-MoVA-36B-A4B-Q4_K_M.gguf Q4_K_M 22.4 GB 24 GB+ (RTX 3090/4090) 32 GB Recommended balance of quality and memory efficiency.
K2-Horizon-MoVA-36B-A4B-Q5_K_M.gguf Q5_K_M 26.4 GB 32 GB 36 GB Near-lossless retention; high precision.
K2-Horizon-MoVA-36B-A4B-Q6_K.gguf Q6_K 30.8 GB 36 GB 40 GB High precision; preserves subtle router logits.
K2-Horizon-MoVA-36B-A4B-Q8_0.gguf Q8_0 39.8 GB 48 GB (2× 24 GB / A6000) 48 GB Maximum precision; near identical to original BF16.

Installation & Setup

1. Build llama.cpp with K2-Horizon Support

# Clone the official architecture fork
git clone -b model/K2Horizon [https://github.com/MBZUAI-IFM/llama.cpp.git](https://github.com/MBZUAI-IFM/llama.cpp.git)
cd llama.cpp

# Build with hardware acceleration (enable CUDA if running on NVIDIA GPUs)
cmake -B build -DGGML_NATIVE=ON -DCMAKE_BUILD_TYPE=Release # Add -DGGML_CUDA=ON for NVIDIA GPU
cmake --build build -j$(nproc) --target llama-cli llama-server

Quickstart Guide

Running via llama-cli

./llama.cpp/build/bin/llama-cli \
  -m ./K2-Horizon-MoVA-36B-A4B-Q4_K_M.gguf \
  -ngl 99 \
  -c 8192 \
  --temp 1.0 \
  --top-p 0.95 \
  -p "<|ifm|im_start|>user\nWrite an efficient Python script for asynchronous data fetching.<|ifm|im_end|>\n<|ifm|im_start|>assistant\n"

Prompt Format & Reasoning Traces

K2-Horizon uses ChatML-style delimiters with native thinking tags:

<|ifm|im_start|>system
You are a helpful assistant.<|ifm|im_end|>
<|ifm|im_start|>user
Your query goes here.<|ifm|im_end|>
<|ifm|im_start|>assistant
<ifm|think>
[Model generates intermediate reasoning steps here]
</ifm|think>
[Final output generated here]<|ifm|im_end|>