drbaph/SoulX-Singer

认证创作者 drbaph 已认证
🤗 Hugging Face 来源text-to-speechapache-2.011 GBother✓ 20 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo drbaph/SoulX-Singer ./model-folder
需要做种者 →

ComfyUI Custom Node

This repository includes a custom node for ComfyUI integration:

🔗 ComfyUI-SoulX-Singer

Use this custom node to integrate SoulX-Singer into your ComfyUI workflows for seamless singing voice synthesis.

SoulX-Singer: Converted .pt model to .safetensors

bf16 + fp32

Audio Samples

Original Audio

Your browser does not support the audio element.

SpongeBob Voice

Your browser does not support the audio element.

Male Voice

Your browser does not support the audio element.
Towards High-Quality Zero-Shot Singing Voice Synthesis


Overview

SoulX-Singer is a high-fidelity, zero-shot singing voice synthesis model that enables users to generate realistic singing voices for unseen singers. It supports melody-conditioned (F0 contour) and score-conditioned (MIDI notes) control for precise pitch, rhythm, and expression.

For more details, please refer to the paper: SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis.


Features

  • Zero-shot synthesis: Generate singing voices for unseen singers without fine-tuning
  • Melody-conditioned control: Use F0 contour for pitch guidance
  • Score-conditioned control: Use MIDI notes for precise musical notation
  • High-fidelity output: Realistic vocal synthesis with natural expression
  • Safetensors format: Optimized model weights in bf16 + fp32 precision

Citation

If you use SoulX-Singer in your research, please cite:

@article{soulxsinger2025,
  title={SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis},
  author={Soul-AILab},
  journal={arXiv preprint arXiv:2602.07803},
  year={2025}
}

License

This project is licensed under the Apache License 2.0.