Abiray/MiniCPM5-2B-GGUF

🤗 Hugging Face 来源text-generationapache-2.0激活 2B11 GBGGUF✓ 6 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Abiray/MiniCPM5-2B-GGUF ./model-folder
需要做种者 →

MiniCPM Tech Report | MiniCPM Wiki (Chinese) | GitHub Repo | UltraData | Online Demo

English | 中文

Highlights

This repository provides quantized GGUF weights for MiniCPM5-2B, the second model in OpenBMB's MiniCPM5 series. It is a dense 2B Transformer scaling up the proven recipe for on-device deployment, edge AI, and local inference via llama.cpp, Ollama, and LM Studio.

🏆 2B-class open-source SOTA: MiniCPM5-2B achieves state-of-the-art performance against models of similar size and remains highly competitive with 4B-class architectures across code generation, mathematics, 128k long-context comprehension, tool use, and multi-step agentic workflows.


Available GGUF Files

Quantization File Name Size Recommendation / Use Case
Q3_K_M MiniCPM5-2B-Q3_K_M.gguf 1.29 GB Ultra-compact; suitable for tight VRAM or RAM constraints.
Q4_K_S MiniCPM5-2B-Q4_K_S.gguf 1.50 GB Fast 4-bit quantization with minimal memory overhead.
Q4_K_M MiniCPM5-2B-Q4_K_M.gguf 1.56 GB Recommended: Best balance of speed, perplexity, and footprint.
Q5_K_M MiniCPM5-2B-Q5_K_M.gguf 1.81 GB High accuracy; preserves subtle reasoning and code logic.
Q6_K MiniCPM5-2B-Q6_K.gguf 2.07 GB High-fidelity 6-bit quantization; near-identical output to BF16.
Q8_0 MiniCPM5-2B-Q8_0.gguf 2.68 GB Near-lossless 8-bit quantization for maximal benchmark fidelity.

Quickstart Guide

llama.cpp

Run inference using llama-cli:

llama-cli \
  -m MiniCPM5-2B-Q4_K_M.gguf \
  -p "Who are you? Please briefly introduce yourself." \
  -n 256 \
  -c 4096 \
  --temp 1.0 \
  --top-p 0.95