This model, QA-MDT, allows for easy setup and usage for generating music from text prompts. It incorporates a quality-aware training strategy to improve the fidelity of generated music.
A Hugging Face Diffusers implementation is available at
this model and
this space. For more detailed instructions and the official PyTorch implementation, please refer to the project's
Github repository and
project page.