The Aria Dataset is a audio-lyrics alignment dataset that contains the word level annotations for 24 Italian Arias. It also contains the phonetic pronunciations as per the Italian language annotated using the SPARSAR system.
The dataset consists of 24 arias with a total audio time of 1 hour and 53 minutes. The mean duration of the audio in the dataset is 284 seconds while the minimum and maximum durations are 125 and 570 seconds respectively.
The dataset has been… See the full description on the dataset page:
https://huggingface.co/datasets/pushkarjajoria/aria-dataset.