Abiray/K2-Horizon-MoVA-36B-A4B-GGUF

🤗 Hugging Face 来源text-generationapache-2.0激活 4B175 GBGGUF✓ 7 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Abiray/K2-Horizon-MoVA-36B-A4B-GGUF ./model-folder
需要做种者 →

K2-Horizon-MoVA-36B-A4B-GGUF

This repository contains GGUF quantizations of IFM/K2-Horizon-MoVA-36B-A4B, a sparse Mixture-of-Experts model utilizing Mixture-of-Values Attention (MoVA). It features 36B total parameters with only ~4B activated per token, enabling near-small-model inference speeds with 36B-scale intelligence.

[!IMPORTANT] Architecture Support Required: As of early 2026, k2-horizon is a newly introduced architecture not yet merged into upstream mainline llama.cpp. To run or quantize these GGUFs, you must build from the official MBZUAI-IFM fork (model/K2Horizon branch). Running on vanilla llama.cpp will result in unknown model architecture: 'k2-horizon'.


Quantization Overview & Hardware Recommendations

File Name Quant Size Recommended VRAM (Full Offload) Minimum System RAM (CPU) Description
K2-Horizon-MoVA-36B-A4B-Q3_K_S.gguf Q3_K_S 16.4 GB 20 GB (RTX 3090/4090) 24 GB Smallest footprint; lowest memory usage.
K2-Horizon-MoVA-36B-A4B-Q3_K_M.gguf Q3_K_M 17.7 GB 22 GB (RTX 3090/4090) 28 GB Balanced 3-bit quantization.
K2-Horizon-MoVA-36B-A4B-Q4_K_S.gguf Q4_K_S 21.4 GB 24 GB (RTX 3090/4090) 32 GB Great performance/size ratio for 24 GB GPUs.
K2-Horizon-MoVA-36B-A4B-Q4_K_M.gguf Q4_K_M 22.4 GB 24 GB+ (RTX 3090/4090) 32 GB Recommended balance of quality and memory efficiency.
K2-Horizon-MoVA-36B-A4B-Q5_K_M.gguf Q5_K_M 26.4 GB 32 GB 36 GB Near-lossless retention; high precision.
K2-Horizon-MoVA-36B-A4B-Q6_K.gguf Q6_K 30.8 GB 36 GB 40 GB High precision; preserves subtle router logits.
K2-Horizon-MoVA-36B-A4B-Q8_0.gguf Q8_0 39.8 GB 48 GB (2× 24 GB / A6000) 48 GB Maximum precision; near identical to original BF16.

Installation & Setup

1. Build llama.cpp with K2-Horizon Support

# Clone the official architecture fork
git clone -b model/K2Horizon [https://github.com/MBZUAI-IFM/llama.cpp.git](https://github.com/MBZUAI-IFM/llama.cpp.git)
cd llama.cpp

# Build with hardware acceleration (enable CUDA if running on NVIDIA GPUs)
cmake -B build -DGGML_NATIVE=ON -DCMAKE_BUILD_TYPE=Release # Add -DGGML_CUDA=ON for NVIDIA GPU
cmake --build build -j$(nproc) --target llama-cli llama-server

Quickstart Guide

Running via llama-cli

./llama.cpp/build/bin/llama-cli \
  -m ./K2-Horizon-MoVA-36B-A4B-Q4_K_M.gguf \
  -ngl 99 \
  -c 8192 \
  --temp 1.0 \
  --top-p 0.95 \
  -p "<|ifm|im_start|>user\nWrite an efficient Python script for asynchronous data fetching.<|ifm|im_end|>\n<|ifm|im_start|>assistant\n"

Prompt Format & Reasoning Traces

K2-Horizon uses ChatML-style delimiters with native thinking tags:

<|ifm|im_start|>system
You are a helpful assistant.<|ifm|im_end|>
<|ifm|im_start|>user
Your query goes here.<|ifm|im_end|>
<|ifm|im_start|>assistant
<ifm|think>
[Model generates intermediate reasoning steps here]
</ifm|think>
[Final output generated here]<|ifm|im_end|>