AudioX-IFcaps (Instruction-Following) is a large-scale, high-quality multimodal dataset designed for training unified audio and music generation models. The dataset contains over 7 million samples with fine-grained, structured annotations that enable precise control over audio generation, including sound event categories, counts, temporal ordering, and timestamps.
General Audio:… See the full description on the dataset page:
https://huggingface.co/datasets/HKUSTAudio/AudioX-IFcaps.