MiniCPM5-2B ยท Q8_0 browser shards

A lossless GGUF split of OpenBMB's official MiniCPM5-2B-Q8_0.gguf. No re-quantization, pruning, conversion of tensor values, or changes to the model architecture. All 381 tensors were compared by SHA-256 against the original. See verification.json.

Split with llama-gguf-split --split-max-size 512M into six files, each under the browser's 2 GB ArrayBuffer limit. Pass the first shard URL to Wllama 3.6.1; it discovers and downloads the other five in parallel.

Chat and create files entirely in your browser: MiniCPM5 WebGPU

Model by OpenBMB. Browser experience by ProCreations. Original weights are Apache 2.0. Please refer to the upstream model card for model capabilities, training, and limitations.

Downloads last month
18
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ProCreations/MiniCPM5-2B-Q8_0-WebGPU

Quantized
(1)
this model

Space using ProCreations/MiniCPM5-2B-Q8_0-WebGPU 1