MiniCPM5-2B · Q8_0 browser shards
A lossless GGUF split of OpenBMB's official MiniCPM5-2B-Q8_0.gguf.
No re-quantization, pruning, conversion of tensor values, or changes to the model architecture.
All 381 tensors were compared by SHA-256 against the original. See verification.json.
Split with llama-gguf-split --split-max-size 512M into six files, each under the browser's 2 GB ArrayBuffer limit.
Pass the first shard URL to Wllama 3.6.1; it discovers and downloads the other five in parallel.
Chat and create files entirely in your browser: MiniCPM5 WebGPU
Model by OpenBMB. Browser experience by ProCreations. Original weights are Apache 2.0. Please refer to the upstream model card for model capabilities, training, and limitations.