Kimi K3 — 8-Layer Q3_K Serving Slice (5.26 GB)
An 8-layer prefix slice of the merged Kimi K3 model, quantized to a Q3_K
trunk with Q8_0 embeddings/output. Runs on a 2019 laptop (Intel Core
i5-8265U, 7.4 GB single-channel RAM) at ~2.5–2.6 tokens/s single-stream.
Derived from: kimi-k3-merged-2exp.gguf (MergeMoE merge of Kimi K3).
Author: Haitham al-Waisy · Contact: haithamalwaisy@gmail.com · X/Twitter · YouTube
Code: https://github.com/Haitham-alwaisy/Kimi-K3-MergeMoE
Verified behaviors
- Loads cleanly, generates real tokens, exits cleanly
- ~2.5–2.6 tokens/s single-stream decode, ~5.3 t/s aggregate (6 batched streams)
- 218 tensors, 4.90 GiB (5.26 GB decimal)
Usage
hf download HaithamalWaisy/kimi-k3-slice-Q3K-8L.gguf kimi-k3-slice-Q3K-8L.gguf --local-dir .
Caveats
- Not benchmarked for quality end-to-end (see paper Section 4.1)
- Merge erases expert routing; slice keeps only 8 of 93 layers
License
Apache 2.0 (inherited from Kimi K3).