AMAImedia/K2-Horizon-7B-GGUF

🤗 Hugging Face sourcetext-generationapache-2.07B activated155 GBGGUF✓ 30 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo AMAImedia/K2-Horizon-7B-GGUF ./model-folder
Needs a seeder →

⚡ Each donation funds the next large quant.

I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.

🎉 Boosty🦄  |  ☕ Buy Me a Coffee🦄  |  ⭐ DonationAlerts🦄

💚 Thanks to Hugging Face for extra storage.🦄


K2-Horizon-7B — GGUF Quantizations

GGUF quantizations of IFM/K2-Horizon-7B, a dense 7B-parameter causal decoder (K2HorizonForCausalLM, model_type: k2_horizon).

Quantized by NANI-Nithin using a custom pipeline built on the MBZUAI-IFM llama.cpp fork (branch model/K2Horizon).


Model Details

Property Value
Base model IFM/K2-Horizon-7B
Architecture K2HorizonForCausalLM (k2_horizon)
Parameters ~7B (dense decoder, no MoE)
Original dtype BF16
BF16 GGUF size ~18.01 GB
Source GGUF IFM/K2-Horizon-7B-GGUF
llama.cpp fork MBZUAI-IFM/llama.cpp @ model/K2Horizon

Note: These GGUFs carry the k2-horizon architecture token and require the MBZUAI-IFM fork (or upstream llama.cpp once support is merged) to run. Vanilla upstream llama.cpp (as of September 2026) does not support K2HorizonForCausalLM.


Included Files

Standard Quantizations

File Bits/Weight Notes
K2-Horizon-7B-BF16.gguf 16 bpw Source quant, full precision
K2-Horizon-7B-Q8_0.gguf 8 bpw Near-lossless, recommended reference
K2-Horizon-7B-Q6_K.gguf 6 bpw Near-lossless K-quant
K2-Horizon-7B-Q5_K_M.gguf 5 bpw Best quality/size in the 5-bit range
K2-Horizon-7B-Q5_K_S.gguf 5 bpw Smaller 5-bit variant
K2-Horizon-7B-Q5_1.gguf 5 bpw Legacy 5-bit
K2-Horizon-7B-Q5_0.gguf 5 bpw Legacy 5-bit
K2-Horizon-7B-Q4_K_M.gguf 4 bpw Recommended general use
K2-Horizon-7B-Q4_K_S.gguf 4 bpw Smaller 4-bit K-quant
K2-Horizon-7B-Q4_1.gguf 4 bpw Legacy 4-bit
K2-Horizon-7B-Q4_0.gguf 4 bpw Legacy 4-bit
K2-Horizon-7B-Q3_K_L.gguf 3 bpw Large 3-bit K-quant
K2-Horizon-7B-Q3_K_M.gguf 3 bpw Medium 3-bit K-quant
K2-Horizon-7B-Q3_K_S.gguf 3 bpw Small 3-bit K-quant
K2-Horizon-7B-Q2_K.gguf 2 bpw Aggressive compression
K2-Horizon-7B-Q2_K_S.gguf 2 bpw Smaller 2-bit K-quant (imatrix-guided)
K2-Horizon-7B-Q2_0.gguf 2.25 bpw Group-64 2-bit
K2-Horizon-7B-Q1_0.gguf 1.125 bpw Maximum compression

IQ (Importance-Matrix) Quantizations

File Bits/Weight
K2-Horizon-7B-IQ4_NL.gguf ~4 bpw
K2-Horizon-7B-IQ4_XS.gguf ~4 bpw
K2-Horizon-7B-IQ3_M.gguf ~3 bpw
K2-Horizon-7B-IQ3_S.gguf ~3 bpw
K2-Horizon-7B-IQ3_XS.gguf ~3 bpw
K2-Horizon-7B-IQ3_XXS.gguf ~3 bpw
K2-Horizon-7B-IQ2_M.gguf ~2 bpw
K2-Horizon-7B-IQ2_S.gguf ~2 bpw
K2-Horizon-7B-IQ2_XS.gguf ~2 bpw
K2-Horizon-7B-IQ2_XXS.gguf ~2 bpw
K2-Horizon-7B-IQ1_M.gguf 1.75 bpw
K2-Horizon-7B-IQ1_S.gguf 1.56 bpw

Quantization Method

  • Source: IFM's official BF16 GGUF (K2-Horizon-7B-BF16.gguf).
  • imatrix: Computed from Salesforce/wikitext (wikitext-2-raw-v1, 500 rows) with 12 GPU layers offloaded on an RTX 4060 Laptop (8 GB VRAM) due to the 18 GB model size. Applied to all K-quants below Q6 and all IQ quants.
  • Fork: MBZUAI-IFM/llama.cpp, branch model/K2Horizon.

Usage

Requires the MBZUAI-IFM llama.cpp fork (model/K2Horizon branch).

git clone -b model/K2Horizon https://github.com/MBZUAI-IFM/llama.cpp
cd llama.cpp && cmake -B build -DGGML_CUDA=ON && cmake --build build --config Release

./build/bin/llama-cli \
  -m K2-Horizon-7B-Q4_K_M.gguf \
  -p "Hello, I am" \
  -n 128 \
  -ngl 35

License

Weights are released under the same license as the original IFM/K2-Horizon-7B model. Please refer to the original repository for full license terms.


Credits