Sharp-MiniCPM5-2B-GGUF
Dynamic imatrix quants of openbmb/MiniCPM5-2B, a 2.5B dense 131k-context model, carrying our cyber-and-coding-weighted importance matrix, an unsloth-dynamic-style per-tensor quant policy, and the Sharp-MiniCPM chat template with fixes and an improved system prompt.
Which quant should I download?
We recommend the Q6_K_XL for anyone who can run it. Practically lossless, at under half the size of the F16 model, and scoring higher than stock Q8_0 on SWE-bench-Live.
Memory math
99% KL is the 99th percentile of KL divergence against the BF16 weights on held-out code — the rare token where a quant actually diverges. top-1 is how often the quant picks the same next token as BF16. The last three columns are the largest q8_0 KV context that fits beside the weights
on a 4 / 6 / 8 GB card.
| tier | size | 99% KL | top-1 | 4 GB | 6 GB | 8 GB |
|---|---|---|---|---|---|---|
| Q4_K_S | 1.52 GB | 0.333 | 92.3% | 97K | 131K* | 131K* |
| Q4_K_XL | 1.60 GB | 0.274 | 93.0% | 94K | 131K* | 131K* |
| Q5_K_XL | 1.89 GB | 0.086 | 96.3% | 81K | 131K* | 131K* |
| Q6_K_XL | 2.20 GB | 0.024 | 98.1% | 68K | 131K* | 131K* |
| Q8_K_M | 2.93 GB | 0.004 | 99.2% | 36K | 130K | 131K* |
* capped by the model's own 131072 native context, not by VRAM. Context assumes
-ctk q8_0 -ctv q8_0, all layers on the GPU, and 0.5 GiB reserved for compute buffers and runtime;
figures are rounded down to the thousand. Run f16 KV instead and roughly halve them.
On 6 GB and up, the model's own 131K window binds before VRAM does at every tier, so choose on fidelity: Q6_K_XL at 0.024 is near-lossless for 2.20 GB. On 4 GB, Q4_K_S gives you 97K of context, and even the floor tier still agrees with BF16 on 92% of next tokens.
Run it
MiniCPM5-2B is a plain llama architecture, so any recent stock llama.cpp runs it:
llama-server -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q6_K_XL --jinja -ngl 99 \
-c 131072 -ctk q8_0 -ctv q8_0 -fa on \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
- Swap the tag after the colon for another tier (
Q4_K_S,Q4_K_XL,Q5_K_XL,Q8_K_M). --jinjauses the Sharp-MiniCPM template embedded in every file; tool calls come back as standard OpenAItool_calls.- Sampling is openbmb's recommendation (temperature 1.0, top-p 0.95, min-p 0) plus top-k 20, the exact
settings the SWE-bench-Live board above ran with. Keep
--min-p 0.0: llama.cpp's default of 0.05 can trap this model in repetition loops. -ctk q8_0 -ctv q8_0is the KV cache the memory table assumes. With VRAM to spare, drop both flags for anf16cache (about twice the memory per token).
This model takes quantization unusually well
Measured against the same corpus, the same 128 KB held-out slice and the same method as our Sharp-Spark-X2.5-4B ladder:
| tier | MiniCPM5-2B 99% KL | Spark-X2.5-4B 99% KL |
|---|---|---|
| Q4_K_S | 0.333 | 1.156 |
| Q4_K_XL | 0.274 | 0.873 |
| Q5_K_XL | 0.086 | 0.363 |
| Q6_K_XL | 0.024 | 0.178 |
| Q8_K_M | 0.004 | 0.023 |
MiniCPM5-2B loses 3–7× less to quantization at every rung, despite being the smaller model —
the opposite of the usual rule that smaller models have less redundancy to spare. Part of that is
this architecture giving the recipe more to work with: untied embeddings let the output head stay
high while the input lookup drops, and split q/k/v lets the cheap KV projections be raised almost
for free — neither of which Spark's tied head and fused attn_qkv allow.
Read this as quantization robustness, not model quality. Each KL is measured against that model's own BF16, so it says how much the quant damaged the model relative to itself. It says nothing about which model is better at anything.
Architecture notes that shape the ladder
MiniCPM5-2B is a plain llama architecture — stock llama.cpp converts and runs it, no custom build.
Three facts drive the per-tensor policy:
- Embeddings are untied.
output.weightandtoken_embd.weightare separate tensors of 10.6% each.output.weightis the output projection and is pinned high on every tier;token_embdis an input lookup and drops lower, which is where most of the saving at the low tiers comes from. - No sliding-window attention — all 42 layers are full attention, so the bump family is the head and tail blocks rather than a structural minority of layers.
- GQA is 8:1 (16 heads, 2 KV heads, head_dim 128), so
attn_kandattn_vare 0.9% of params each. Every tier raises both; it costs ~11 MB across all 42 layers.
KV cache is 42 KB/token at f16, 22.3 KB/token at q8_0 — larger per token than a bigger model
with windowed attention would need, because every layer here holds a full-length cache.
The Sharp-MiniCPM template
MiniCPM5's own template with six defects repaired and the house terseness prompt spliced in. It is
not a port of another model's template: the control tokens, the <think> convention and the XML
<function>/<param> tool-call format are all MiniCPM5's, untouched, because the model was trained
on them. Every fix was reproduced against the stock template before being written.
- Assistant text after a
<tool_sep>was silently discarded. Stock builds the interleaved content with{% set %}inside a{% for %}; Jinja throws those assignments away when the loop ends, so the content reverted to the text before the first separator and everything after it vanished. Rebuilt on a namespace. - Content blocks were dropped to an empty string — a user turn sent as
[{"type": "text", ...}]arrived completely blank. - Tool results given as content blocks were dumped into the prompt as raw JSON.
- Non-string tool-call arguments rendered as a Python repr — a nested object reached the model
as
{'a': 1, 'b': True}, single quotes and all, instead of JSON. - Stale reasoning was replayed for every historical turn. Stock computes
ns.last_query_indexfor exactly this purpose and then never reads it; we wired it up. - CDATA wrapping did not escape a literal
]]>, which closes the block early and corrupts the rest of the call.
Additions: a terseness instruction force-appended to the system prompt (opt out with
chat_template_kwargs={"terse": false}), and suppress_tool_instructions, which stands the tool
block down when the runtime has injected its own tool protocol.
Verified to render in both jinja2 (transformers) and minja (llama.cpp). It does not render in
minijinja, whose tojson rejects the ensure_ascii keyword — that construct is inherited from the
stock template and is kept deliberately, because dropping it would \u-escape CJK in tool
descriptions and argument values on a bilingual model.
The template ships in this repo as chat_template.jinja.