ddalcu/Kokoro-82M-MLX-Serve

🤗 Hugging Face 来源text-to-speechapache-2.081M 参数325 MBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo ddalcu/Kokoro-82M-MLX-Serve ./model-folder
需要做种者 →

Kokoro-82M for mlx-serve

Kokoro-82M repacked for mlx-serve's native Zig + MLX engine.

Same weights, same precision — f32, not quantized. The repack folds weight-norm into .weight (so the engine has no weight_g/weight_v to resolve), drops the unused pooler tensors, converts .pth to safetensors, and bundles the 54 voice packs plus the English pronunciation dictionaries the phonemizer needs.

File Contents
model.safetensors 457 tensors, 81.1M params, 325 MB f32
voices.safetensors 54 voice packs, [510, 1, 256] each
g2p/*.json misaki us_gold / us_silver / gb_gold

Verified against the torch reference: per-phoneme durations match exactly, F0 / noise / text-encoder outputs at cosine 1.000000, and the waveform at 0.9968 — which is inside the reference's own seed-to-seed spread (0.9941–0.9960), since its vocoder is stochastic.

About 17x realtime on an M-series Mac, ~350 MB resident.

Run it

Download MLX Core.app, open Settings ▸ Voice, and pick Kokoro as the voice engine. 54 voices, and naming several separated by commas blends them into a new one.

Over HTTP:

curl -X POST http://localhost:11234/v1/audio/speech \
  -H 'content-type: application/json' \
  -d '{"model":"kokoro","input":"Hello there.","voice":"af_bella,af_sky"}' \
  --output out.wav

Credit

Kokoro-82M and the misaki G2P dictionaries are both by hexgrad, Apache-2.0. This repo only changes the packaging. No espeak-ng anywhere in the pipeline — the dictionaries make it unnecessary, which keeps the whole path Apache-2.0.