aufklarer/VoxCPM2-MLX-bf16

🤗 Hugging Face 来源text-to-speechapache-2.02.4B 参数5.0 GBsafetensors✓ 3 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo aufklarer/VoxCPM2-MLX-bf16 ./model-folder
需要做种者 →

VoxCPM2 — MLX bf16

Full-precision MLX port for Apple Silicon.

MLX port of openbmb/VoxCPM2 — a 2B-parameter multilingual diffusion-autoregressive TTS model with 48 kHz studio-quality output, voice cloning, and instruction-driven voice design.

Part of soniqo.audio — an on-device speech toolkit for Apple Silicon. Consumed by the open-source speech-swift library (module VoxCPM2TTS).

Bundle size: 4.96 GB

Use cases

Variants

Variant Size Notes
bf16 ~5.0 GB Reference quality, no Linear quantization.
int8 ~3.0 GB 8-bit group quantization. Mean rel-L2 0.53 % vs bf16.

Capabilities

  • 30 languages including English, Chinese, Indonesian, Japanese, Korean
  • 48 kHz output
  • Zero-shot synthesis — generate speech from text alone
  • Voice cloning — clone a target speaker from a single reference clip
  • Voice design — natural-language style control (e.g. "young female voice, warm and gentle")
  • Ultimate cloning — reference audio + transcript for prosody-preserving cloning
  • Streaming generation — patch-level decoding for low-latency synthesis

Precision

No quantization. All Linear weights stored as bfloat16. Use this variant for reference quality or when memory is not a constraint.

Usage with speech-swift

This bundle is consumed by soniqo/speech-swift's VoxCPM2TTS Swift module.

import VoxCPM2TTS

let model = try await VoxCPM2TTSModel.fromPretrained(
    modelId: "aufklarer/VoxCPM2-MLX-bf16"
)
let audio = try await model.generate(text: "Hello from VoxCPM2.", language: "english")

Or via the CLI:

speech speak "Hello from VoxCPM2." --engine voxcpm2 --voxcpm2-variant bf16 -o hi.wav

Source

This bundle is converted from the upstream PyTorch weights at openbmb/VoxCPM2.

License

Apache 2.0 — inherited from the upstream openbmb/VoxCPM2 model.

Responsible use

Voice cloning capability is included. Users are responsible for obtaining consent for any voice that is cloned and for not using the model to impersonate individuals without their permission, generate disinformation, or commit fraud.