WhiskyAKM/MiniCPM5-2B-GGUF

🤗 Hugging Face sourcetext-generationapache-2.02B activated19 GBGGUF✓ 11 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo WhiskyAKM/MiniCPM5-2B-GGUF ./model-folder
Needs a seeder →

MiniCPM5-2B GGUF

GGUF quantized version of openbmb/MiniCPM5-2B, the second model in the MiniCPM5 series. It is a dense 2B Transformer (LlamaForCausalLM) built for on-device, local deployment, and resource-constrained scenarios, reaching 2B-class open-source SOTA.

Model Overview

MiniCPM5-2B is designed for local assistants, coding agents, tool-use workflows, and reasoning scenarios where a compact model is preferred. It keeps a small deployment footprint while providing native long-context support (131,072 tokens), and was post-trained with RL + OPD (On-Policy Distillation). It supports a chat template with thinking/reasoning mode (controlled via enable_thinking) and XML-style tool calling.

Model Architecture

Property Value
Architecture LlamaForCausalLM (dense)
Total Parameters 2.52B
Non-Embedding Parameters 1.98B
Layers 42
Attention Heads (GQA) 16 Q / 2 KV
Context Length 131,072 (128K)
Original Precision bfloat16
Supported Languages en, zh

Available GGUF Files

File Quantization Size Use Case
minicpm5-2b-bf16.gguf BF16 4.7 GB Full precision (reference quality)
minicpm5-2b-Q8_0.gguf Q8_0 2.5 GB Near-lossless, low overhead
minicpm5-2b-Q6_K.gguf Q6_K 2.0 GB Very high quality
minicpm5-2b-Q5_K_M.gguf Q5_K_M 1.7 GB High quality, balanced
minicpm5-2b-Q5_K_S.gguf Q5_K_S 1.7 GB High quality, slightly smaller
minicpm5-2b-Q4_K_M.gguf Q4_K_M 1.5 GB Recommended: good quality/size trade-off
minicpm5-2b-Q4_K_S.gguf Q4_K_S 1.4 GB Smaller, slight quality loss
minicpm5-2b-Q4_0.gguf Q4_0 1.4 GB Legacy 4-bit, broad compatibility

DSpark Draft Model (Speculative Decoding)

File Quantization Size Use Case
minicpm5-2b-dspark-bf16.gguf BF16 623 MB Draft model, full precision
minicpm5-2b-dspark-Q8_0.gguf Q8_0 334 MB Draft model, near-lossless
minicpm5-2b-dspark-Q4_K_M.gguf Q4_K_M 186 MB Draft model, smallest footprint

These are the MiniCPM5-2B-DSpark draft models, trained for speculative decoding with MiniCPM5-2B. They accelerate generation without changing the target model's outputs.

Usage

llama.cpp CLI

./llama-cli \
  -m minicpm5-2b-Q4_K_M.gguf \
  -p "Explain quantum computing in simple terms." \
  --temp 1.0 --top-p 0.95

llama-server (OpenAI-compatible API)

./llama-server \
  -m minicpm5-2b-Q4_K_M.gguf \
  --host 0.0.0.0 --port 8080

The GGUF also works with Ollama and LM Studio.

Speculative Decoding (DSpark)

Pair the target model with a DSpark draft model to speed up inference. Draft model quality has minimal impact on output, so smaller quants (e.g. Q4_K_M) are usually fine:

./llama-server \
  -m minicpm5-2b-Q4_K_M.gguf \
  -md minicpm5-2b-dspark-Q4_K_M.gguf \
  --host 0.0.0.0 --port 8080

Thinking Mode

The model supports deep-thinking output. You can control it per request via the chat template, e.g. with an OpenAI-compatible API:

"chat_template_kwargs": {"enable_thinking": false}

Tool Calling

MiniCPM5-2B emits XML-style tool calls. Tool definitions are injected into the prompt, and tool results are returned in the observation/tool role. SGLang's built-in minicpm5 parser converts these to OpenAI-compatible tool_calls natively (see upstream model card).

Generation Parameters

Recommended parameters from the original model:

Parameter Value
Temperature 1.0
Top-P 0.95

Quantization

These GGUF files were created from the BF16 source model using llama-quantize from the llama.cpp project.

Acknowledgements

License

Apache-2.0 License