Sharp-MiniCPM5-2B-MLX-oQ6e
An Apple-silicon MLX build of openbmb/MiniCPM5-2B, a 2.5B dense 131k-context model, re-quantized with oMLX's oQ6e quantizer against our own code- and cyber-weighted importance-matrix corpus, and carrying the Sharp-MiniCPM chat template with fixes and an improved system prompt inside the checkpoint.
This is the MLX sibling of the GGUF ladder. Same base weights, same calibration corpus, same chat template — a different quantizer (oMLX's oQ instead of llama.cpp's k-quants), for people who run models through MLX on a Mac.
The numbers below were measured on the GGUF build, not this one
The board below is the best evidence we have about this model, and we would rather show it than show
nothing — but it was produced on the GGUF build
at its Q6_K_XL tier, using llama.cpp. This file uses a different quantizer (oMLX's oQ), and changing
quantizer moves results. Read the board as evidence about the model, not as a measurement of this
file. If you need numbers you can hold us to, use the GGUF build.
SWE-bench-Live is an autonomous agentic software-engineering benchmark: it tests the model's ability to solve a set of real issues and bugs in open-source codebases, published continuously and recently, with hidden regression tests that catch if the model breaks something trying to fix something. Doing well on it represents real-world autonomous coding ability — the opposite of benchmark memorization.
Run it
This is a text model (llama architecture), so load it with mlx-lm, not mlx-vlm. At
≈2.1 GB of unified memory it runs on any Apple-silicon Mac with room to spare, leaving the whole
131k context available on 8 GB and up.
mlx-lm:
hf download peculiar-ragdoll/Sharp-MiniCPM5-2B-MLX-oQ6e --local-dir SharpMiniCPM-MLX
python -m mlx_lm generate --model SharpMiniCPM-MLX --max-tokens 512 \
--prompt "Write a Python function that returns the nth Fibonacci number."
Or serve an OpenAI-compatible endpoint:
python -m mlx_lm server --model SharpMiniCPM-MLX --port 8080
oMLX — put the folder under ~/.omlx/models/peculiar-ragdoll/Sharp-MiniCPM5-2B-MLX-oQ6e, or pull
it from the oMLX model manager, and it serves with the embedded Sharp-MiniCPM template and tool calling.
- Sampling: openbmb's recommendation — temperature 1.0, top-p 0.95, min-p 0 — plus top-k 20, the settings the SWE-bench-Live board above ran with.
- Budget:
mlx_lm generateneeds an explicit--max-tokens; the512above is sized for a one-shot demo, not real work. Give real work a generous ceiling (32768if you cap it at all). A low budget degrades performance and will not make the model converge any faster. - Template: the Sharp-MiniCPM template is embedded, so tool calls come back in MiniCPM5's native format. See the template section below.
The oQ6e imatrix — calibrated on our corpus
oMLX's oQ quantizer runs its own importance-matrix pass — the "e" in oQ6e — that measures which
weights carry the most signal before deciding what to keep at higher precision. For this build we fed
that pass our own code- and cybersecurity-weighted calibration corpus — the same corpus behind the
GGUF build's importance matrix —
rather than oMLX's default calibration set.
An imatrix is not training data. It measures which weights carry the load under a representative input distribution, so the quantizer spends its precision there and lets rounding error fall where it matters least. Point that measurement at code and the quant stays close to full precision on exactly the work this build is for.
oQ derives its own importance data from that corpus — it does not consume the GGUF imatrix bytes we baked. Same corpus, different quantizer and a different importance computation, so this build is not identical to any GGUF tier. Only the GGUF route has been benchmarked.
oQ6e is oMLX's dynamic quantizer: a 6-bit base with mixed precision by layer position and selective non-quantization, plus the imatrix pass above. It is the same idea as the GGUF ladder's dynamic quants, computed by a different tool.
Architecture notes that shape the quant
MiniCPM5-2B is a plain llama architecture. Three facts drive the per-tensor policy:
- Embeddings are untied.
output.weightandtoken_embd.weightare separate tensors of 10.6% each. The output projection is pinned high; the input lookup drops lower, which is where most of the low-bit saving comes from. - No sliding-window attention — all 42 layers are full attention, so the higher-precision family is the head and tail blocks rather than a structural minority of layers.
- GQA is 8:1 (16 heads, 2 KV heads, head_dim 128), so
attn_kandattn_vare cheap to raise and both are kept high.
The Sharp-MiniCPM template
MiniCPM5's own template with six defects repaired and the house terseness prompt spliced in. It is not
a port of another model's template: the control tokens, the <think> convention and the XML
<function>/<param> tool-call format are all MiniCPM5's, untouched, because the model was trained on
them. Every fix was reproduced against the stock template before being written.
- Assistant text after a
<tool_sep>was silently discarded. Stock builds the interleaved content with{% set %}inside a{% for %}; Jinja throws those assignments away when the loop ends, so the content reverted to the text before the first separator and everything after it vanished. Rebuilt on a namespace. - Content blocks were dropped to an empty string — a user turn sent as
[{"type": "text", ...}]arrived completely blank. - Tool results given as content blocks were dumped into the prompt as raw JSON.
- Non-string tool-call arguments rendered as a Python repr — a nested object reached the model as
{'a': 1, 'b': True}, single quotes and all, instead of JSON. - Stale reasoning was replayed for every historical turn. Stock computes
ns.last_query_indexfor exactly this purpose and then never reads it; we wired it up. - CDATA wrapping did not escape a literal
]]>, which closes the block early and corrupts the rest of the call.
Additions: a terseness instruction force-appended to the system prompt (opt out with
chat_template_kwargs={"terse": false}), and suppress_tool_instructions, which stands the tool block
down when the runtime has injected its own tool protocol.
The template ships in this repo as chat_template.jinja.
Links
- Sharp-MiniCPM5-2B GGUF ladder — the benchmarked build, with the full quant ladder and KL-divergence fidelity curve.
- openbmb/MiniCPM5-2B — the base model.
- oMLX — the oQ dynamic quantizer this build uses.