peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF

🤗 Hugging Face 来源text-generationapache-2.0激活 2B10 GBGGUF✓ 9 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF ./model-folder
需要做种者 →

Sharp-MiniCPM5-2B-GGUF

Dynamic imatrix quants of openbmb/MiniCPM5-2B, a 2.5B dense 131k-context model, carrying our cyber-and-coding-weighted importance matrix, an unsloth-dynamic-style per-tensor quant policy, and the Sharp-MiniCPM chat template with fixes and an improved system prompt.

Which quant should I download?

We recommend the Q6_K_XL for anyone who can run it. Practically lossless, at under half the size of the F16 model, and scoring higher than stock Q8_0 on SWE-bench-Live.

Memory math

99% KL is the 99th percentile of KL divergence against the BF16 weights on held-out code — the rare token where a quant actually diverges. top-1 is how often the quant picks the same next token as BF16. The last three columns are the largest q8_0 KV context that fits beside the weights on a 4 / 6 / 8 GB card.

tier size 99% KL top-1 4 GB 6 GB 8 GB
Q4_K_S 1.52 GB 0.333 92.3% 97K 131K* 131K*
Q4_K_XL 1.60 GB 0.274 93.0% 94K 131K* 131K*
Q5_K_XL 1.89 GB 0.086 96.3% 81K 131K* 131K*
Q6_K_XL 2.20 GB 0.024 98.1% 68K 131K* 131K*
Q8_K_M 2.93 GB 0.004 99.2% 36K 130K 131K*

* capped by the model's own 131072 native context, not by VRAM. Context assumes -ctk q8_0 -ctv q8_0, all layers on the GPU, and 0.5 GiB reserved for compute buffers and runtime; figures are rounded down to the thousand. Run f16 KV instead and roughly halve them.

On 6 GB and up, the model's own 131K window binds before VRAM does at every tier, so choose on fidelity: Q6_K_XL at 0.024 is near-lossless for 2.20 GB. On 4 GB, Q4_K_S gives you 97K of context, and even the floor tier still agrees with BF16 on 92% of next tokens.

Run it

MiniCPM5-2B is a plain llama architecture, so any recent stock llama.cpp runs it:

llama-server -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q6_K_XL --jinja -ngl 99 \
  -c 131072 -ctk q8_0 -ctv q8_0 -fa on \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
  • Swap the tag after the colon for another tier (Q4_K_S, Q4_K_XL, Q5_K_XL, Q8_K_M).
  • --jinja uses the Sharp-MiniCPM template embedded in every file; tool calls come back as standard OpenAI tool_calls.
  • Sampling is openbmb's recommendation (temperature 1.0, top-p 0.95, min-p 0) plus top-k 20, the exact settings the SWE-bench-Live board above ran with. Keep --min-p 0.0: llama.cpp's default of 0.05 can trap this model in repetition loops.
  • -ctk q8_0 -ctv q8_0 is the KV cache the memory table assumes. With VRAM to spare, drop both flags for an f16 cache (about twice the memory per token).

This model takes quantization unusually well

Measured against the same corpus, the same 128 KB held-out slice and the same method as our Sharp-Spark-X2.5-4B ladder:

tier MiniCPM5-2B 99% KL Spark-X2.5-4B 99% KL
Q4_K_S 0.333 1.156
Q4_K_XL 0.274 0.873
Q5_K_XL 0.086 0.363
Q6_K_XL 0.024 0.178
Q8_K_M 0.004 0.023

MiniCPM5-2B loses 3–7× less to quantization at every rung, despite being the smaller model — the opposite of the usual rule that smaller models have less redundancy to spare. Part of that is this architecture giving the recipe more to work with: untied embeddings let the output head stay high while the input lookup drops, and split q/k/v lets the cheap KV projections be raised almost for free — neither of which Spark's tied head and fused attn_qkv allow.

Read this as quantization robustness, not model quality. Each KL is measured against that model's own BF16, so it says how much the quant damaged the model relative to itself. It says nothing about which model is better at anything.

Architecture notes that shape the ladder

MiniCPM5-2B is a plain llama architecture — stock llama.cpp converts and runs it, no custom build. Three facts drive the per-tensor policy:

  1. Embeddings are untied. output.weight and token_embd.weight are separate tensors of 10.6% each. output.weight is the output projection and is pinned high on every tier; token_embd is an input lookup and drops lower, which is where most of the saving at the low tiers comes from.
  2. No sliding-window attention — all 42 layers are full attention, so the bump family is the head and tail blocks rather than a structural minority of layers.
  3. GQA is 8:1 (16 heads, 2 KV heads, head_dim 128), so attn_k and attn_v are 0.9% of params each. Every tier raises both; it costs ~11 MB across all 42 layers.

KV cache is 42 KB/token at f16, 22.3 KB/token at q8_0 — larger per token than a bigger model with windowed attention would need, because every layer here holds a full-length cache.

The Sharp-MiniCPM template

MiniCPM5's own template with six defects repaired and the house terseness prompt spliced in. It is not a port of another model's template: the control tokens, the <think> convention and the XML <function>/<param> tool-call format are all MiniCPM5's, untouched, because the model was trained on them. Every fix was reproduced against the stock template before being written.

  1. Assistant text after a <tool_sep> was silently discarded. Stock builds the interleaved content with {% set %} inside a {% for %}; Jinja throws those assignments away when the loop ends, so the content reverted to the text before the first separator and everything after it vanished. Rebuilt on a namespace.
  2. Content blocks were dropped to an empty string — a user turn sent as [{"type": "text", ...}] arrived completely blank.
  3. Tool results given as content blocks were dumped into the prompt as raw JSON.
  4. Non-string tool-call arguments rendered as a Python repr — a nested object reached the model as {'a': 1, 'b': True}, single quotes and all, instead of JSON.
  5. Stale reasoning was replayed for every historical turn. Stock computes ns.last_query_index for exactly this purpose and then never reads it; we wired it up.
  6. CDATA wrapping did not escape a literal ]]>, which closes the block early and corrupts the rest of the call.

Additions: a terseness instruction force-appended to the system prompt (opt out with chat_template_kwargs={"terse": false}), and suppress_tool_instructions, which stands the tool block down when the runtime has injected its own tool protocol.

Verified to render in both jinja2 (transformers) and minja (llama.cpp). It does not render in minijinja, whose tojson rejects the ensure_ascii keyword — that construct is inherited from the stock template and is kept deliberately, because dropping it would \u-escape CJK in tool descriptions and argument values on a bilingual model.

The template ships in this repo as chat_template.jinja.