ddalcu/ACE-Step-1.5-XL-Turbo-MLX-Serve-4bit

🤗 Hugging Face 来源text-to-audiomit762M 参数1.5 GBsafetensors✓ 4 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo ddalcu/ACE-Step-1.5-XL-Turbo-MLX-Serve-4bit ./model-folder
需要做种者 →

ACE-Step 1.5 XL Turbo — MLX (4-bit) for mlx-serve

Native Apple Silicon build of ACE-Step 1.5 XL Turbo (4-billion-parameter music-generation DiT, 8-step distilled, no CFG) for mlx-serve. Type a style prompt ("upbeat synthwave with driving bass"), optionally add lyrics, and get an original 48 kHz stereo track — entirely on-device.

Contents (one self-contained bundle, ~4.0 GB)

File What
model.safetensors 32-layer DiT decoder + condition encoder + silence latent. Large linears 4-bit affine (group 64); the timestep-embedding family stays 8-bit — its adaLN scale/shift modulates every layer and few-step turbo models compound modulation error; norms/convs/small projections bf16. Conv layouts pre-swapped to MLX [out, K, in].
vae.safetensors AutoencoderOobleck audio VAE (48 kHz stereo, hop 1920 → 25 Hz latents). Weight-norm fused, bf16 (Snake α/β fp32) — audio VAEs are precision-critical, so no quantization here.
text_encoder/ Qwen3-Embedding-0.6B verbatim (bf16, standard qwen3) — encodes the style prompt; its embedding table encodes lyrics.
config.json {"model_type": "acestep", "quant": "4bit", ...} — the marker mlx-serve's audio engine dispatches on.

No external dependencies — text encoder and VAE ride in the bundle. mlx-serve infers each tensor's (bits, group) from packed geometry, so this mixed 4/8-bit checkpoint loads through the same code path as the 8-bit build.

Use

Server API:

mlx-serve --serve --model-dir ~/.mlx-serve/models
curl -X POST http://127.0.0.1:8080/v1/audio/music-generations \
  -H 'Content-Type: application/json' \
  -d '{"model": "ACE-Step-1.5-XL-Turbo-MLX-Serve-4bit",
       "prompt": "upbeat synthwave with driving bass, dreamy pads",
       "duration_seconds": 30, "seed": 7}' \
  -o track.wav

Fields: prompt (required), lyrics (empty → instrumental), vocal_language, bpm (30–300), keyscale (e.g. "F# minor"), timesignature (2/3/4/6), duration_seconds (10–600), seed, stream (SSE progress: encode → 8 diffusion steps → chunked VAE decode).

Conversion & fidelity

Converted by mlx-serve's tests/convert_acestep_weights.py --bits 4 from the fp32 source checkpoint. The full pipeline (Qwen3 text encoding, condition encoders, the 32-layer DiT, the flow-match sampler with DCW correction, and the Oobleck VAE) is re-implemented natively in Zig on MLX and validated against the fp32 PyTorch reference with cosine-similarity oracles (measured on the 8-bit build; the 4-bit build shares every code path and differs only in weight precision). Same-seed A/B clips against the 8-bit build show matching loudness and spectral balance; expect slightly softer detail than 8-bit, most audible on dense vocal mixes.

License & credits

MIT (see LICENSE). Original model by ACE Studio and StepFun — trained on licensed, royalty-free, and synthetic data; generated music is commercially usable per the upstream project. Text encoder: Qwen3-Embedding-0.6B (Qwen team, Apache 2.0).