This dataset was used to Pre-train Tango-Full-FT-Audiocaps. TangoPromptBank is a diverse corpus consisting of textual prompts and audio samples sourced from WavCaps [1], AudioCaps [9], ESC [2], UrbanSound [3], MusicCaps [4], GTZAN [5], and Musical Instruments [6] dataset. The dataset statistics are reported in Table 1. All audio clips longer than 10 seconds were segmented into partitions of successive 10… See the full description on the dataset page:
https://huggingface.co/datasets/declare-lab/TangoPromptBank.