peculiar-ragdoll/Sharp-MiniCPM5-2B-MLX-oQ6e

🤗 Hugging Face 来源text-generationapache-2.02.5B 参数5.0 GBsafetensors✓ 3 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo peculiar-ragdoll/Sharp-MiniCPM5-2B-MLX-oQ6e ./model-folder
需要做种者 →

Sharp-MiniCPM5-2B-MLX-oQ6e

An Apple-silicon MLX build of openbmb/MiniCPM5-2B, a 2.5B dense 131k-context model, re-quantized with oMLX's oQ6e quantizer against our own code- and cyber-weighted importance-matrix corpus, and carrying the Sharp-MiniCPM chat template with fixes and an improved system prompt inside the checkpoint.

This is the MLX sibling of the GGUF ladder. Same base weights, same calibration corpus, same chat template — a different quantizer (oMLX's oQ instead of llama.cpp's k-quants), for people who run models through MLX on a Mac.

The numbers below were measured on the GGUF build, not this one

The board below is the best evidence we have about this model, and we would rather show it than show nothing — but it was produced on the GGUF build at its Q6_K_XL tier, using llama.cpp. This file uses a different quantizer (oMLX's oQ), and changing quantizer moves results. Read the board as evidence about the model, not as a measurement of this file. If you need numbers you can hold us to, use the GGUF build.

SWE-bench-Live is an autonomous agentic software-engineering benchmark: it tests the model's ability to solve a set of real issues and bugs in open-source codebases, published continuously and recently, with hidden regression tests that catch if the model breaks something trying to fix something. Doing well on it represents real-world autonomous coding ability — the opposite of benchmark memorization.

Run it

This is a text model (llama architecture), so load it with mlx-lm, not mlx-vlm. At ≈2.1 GB of unified memory it runs on any Apple-silicon Mac with room to spare, leaving the whole 131k context available on 8 GB and up.

mlx-lm:

hf download peculiar-ragdoll/Sharp-MiniCPM5-2B-MLX-oQ6e --local-dir SharpMiniCPM-MLX
python -m mlx_lm generate --model SharpMiniCPM-MLX --max-tokens 512 \
  --prompt "Write a Python function that returns the nth Fibonacci number."

Or serve an OpenAI-compatible endpoint:

python -m mlx_lm server --model SharpMiniCPM-MLX --port 8080

oMLX — put the folder under ~/.omlx/models/peculiar-ragdoll/Sharp-MiniCPM5-2B-MLX-oQ6e, or pull it from the oMLX model manager, and it serves with the embedded Sharp-MiniCPM template and tool calling.

  • Sampling: openbmb's recommendation — temperature 1.0, top-p 0.95, min-p 0 — plus top-k 20, the settings the SWE-bench-Live board above ran with.
  • Budget: mlx_lm generate needs an explicit --max-tokens; the 512 above is sized for a one-shot demo, not real work. Give real work a generous ceiling (32768 if you cap it at all). A low budget degrades performance and will not make the model converge any faster.
  • Template: the Sharp-MiniCPM template is embedded, so tool calls come back in MiniCPM5's native format. See the template section below.

The oQ6e imatrix — calibrated on our corpus

oMLX's oQ quantizer runs its own importance-matrix pass — the "e" in oQ6e — that measures which weights carry the most signal before deciding what to keep at higher precision. For this build we fed that pass our own code- and cybersecurity-weighted calibration corpus — the same corpus behind the GGUF build's importance matrix — rather than oMLX's default calibration set.

An imatrix is not training data. It measures which weights carry the load under a representative input distribution, so the quantizer spends its precision there and lets rounding error fall where it matters least. Point that measurement at code and the quant stays close to full precision on exactly the work this build is for.

oQ derives its own importance data from that corpus — it does not consume the GGUF imatrix bytes we baked. Same corpus, different quantizer and a different importance computation, so this build is not identical to any GGUF tier. Only the GGUF route has been benchmarked.

oQ6e is oMLX's dynamic quantizer: a 6-bit base with mixed precision by layer position and selective non-quantization, plus the imatrix pass above. It is the same idea as the GGUF ladder's dynamic quants, computed by a different tool.

Architecture notes that shape the quant

MiniCPM5-2B is a plain llama architecture. Three facts drive the per-tensor policy:

  1. Embeddings are untied. output.weight and token_embd.weight are separate tensors of 10.6% each. The output projection is pinned high; the input lookup drops lower, which is where most of the low-bit saving comes from.
  2. No sliding-window attention — all 42 layers are full attention, so the higher-precision family is the head and tail blocks rather than a structural minority of layers.
  3. GQA is 8:1 (16 heads, 2 KV heads, head_dim 128), so attn_k and attn_v are cheap to raise and both are kept high.

The Sharp-MiniCPM template

MiniCPM5's own template with six defects repaired and the house terseness prompt spliced in. It is not a port of another model's template: the control tokens, the <think> convention and the XML <function>/<param> tool-call format are all MiniCPM5's, untouched, because the model was trained on them. Every fix was reproduced against the stock template before being written.

  1. Assistant text after a <tool_sep> was silently discarded. Stock builds the interleaved content with {% set %} inside a {% for %}; Jinja throws those assignments away when the loop ends, so the content reverted to the text before the first separator and everything after it vanished. Rebuilt on a namespace.
  2. Content blocks were dropped to an empty string — a user turn sent as [{"type": "text", ...}] arrived completely blank.
  3. Tool results given as content blocks were dumped into the prompt as raw JSON.
  4. Non-string tool-call arguments rendered as a Python repr — a nested object reached the model as {'a': 1, 'b': True}, single quotes and all, instead of JSON.
  5. Stale reasoning was replayed for every historical turn. Stock computes ns.last_query_index for exactly this purpose and then never reads it; we wired it up.
  6. CDATA wrapping did not escape a literal ]]>, which closes the block early and corrupts the rest of the call.

Additions: a terseness instruction force-appended to the system prompt (opt out with chat_template_kwargs={"terse": false}), and suppress_tool_instructions, which stands the tool block down when the runtime has injected its own tool protocol.

The template ships in this repo as chat_template.jinja.

Links