HaithamalWaisy/kimi-k3-slice-Q3K-8L.gguf

🤗 On Hugging Faceapache-2.05.3 GBGGUFHF checksums availableupdated today
Magnet

Kimi K3 — 8-Layer Q3_K Serving Slice (5.26 GB)

An 8-layer prefix slice of the merged Kimi K3 model, quantized to a Q3_K

trunk with Q8_0 embeddings/output. Runs on a 2019 laptop (Intel Core

i5-8265U, 7.4 GB single-channel RAM) at ~2.5–2.6 tokens/s single-stream.

Derived from: kimi-k3-merged-2exp.gguf (MergeMoE merge of Kimi K3).

Paper: Kimi K3 Under Compute Constraints: Implementing MergeMoE to Take 14 Shards to 5GB on Consumer Hardware

Author: Haitham al-Waisy · Contact: haithamalwaisy@gmail.com · X/Twitter · YouTube

Code: https://github.com/Haitham-alwaisy/Kimi-K3-MergeMoE

Verified behaviors

  • Loads cleanly, generates real tokens, exits cleanly
  • ~2.5–2.6 tokens/s single-stream decode, ~5.3 t/s aggregate (6 batched streams)
  • 218 tensors, 4.90 GiB (5.26 GB decimal)

Usage

hf download HaithamalWaisy/kimi-k3-slice-Q3K-8L.gguf kimi-k3-slice-Q3K-8L.gguf --local-dir .

Caveats

  • Not benchmarked for quality end-to-end (see paper Section 4.1)
  • Merge erases expert routing; slice keeps only 8 of 93 layers

License

Apache 2.0 (inherited from Kimi K3).