ACE-Step 1.5 XL — Base (4B DiT)
Project |
Hugging Face |
ModelScope |
Space Demo |
Discord |
Tech Report
Model Details
This is the XL (4B) Base variant of ACE-Step 1.5 — a larger DiT decoder with ~4B parameters for higher audio quality. It is the foundation model supporting all tasks: text-to-music, cover, repaint, extract, lego, and complete.
XL Architecture
GPU Requirements
All LM models (0.6B / 1.7B / 4B) are fully compatible with XL.
Key Features
- 💰 Commercial-Ready: Trained on legally compliant datasets. Generated music can be used for commercial purposes.
- 📚 Safe Training Data: Licensed music, royalty-free/public domain, and synthetic (MIDI-to-Audio) data.
- 🎯 Full Task Support: Text2Music, Cover, Repaint, Extract, Lego, Complete.
- 🔮 Higher Quality: 4B parameters provide richer audio quality compared to the 2B variants.
Quick Start
# Install ACE-Step
git clone https://github.com/ace-step/ACE-Step-1.5.git
cd ACE-Step-1.5
pip install -e .
# Download this model
huggingface-cli download ACE-Step/acestep-v15-xl-base --local-dir ./checkpoints/acestep-v15-xl-base
# Run with Gradio UI
python acestep --config-path acestep-v15-xl-base
Model Zoo
XL (4B) DiT Models
2B DiT Models
LM Models (all compatible with XL)
Acknowledgements
This project is co-led by ACE Studio and StepFun.
Citation
@misc{gong2026acestep,
title={ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation},
author={Junmin Gong, Yulin Song, Wenxiao Zhao, Sen Wang, Shengyuan Xu, Jing Guo},
howpublished={\url{https://github.com/ace-step/ACE-Step-1.5}},
year={2026},
note={GitHub repository}
}