One video may contain different audio-visual events, so the total number of videos is not 4143.
annotations.txt: annotations of the AVE dataset. For each sample, you can find its event category, YouTube ID, Quality (all good, meaning that it contains an AVE), start time of an audio-visual event, and the end time of an audio-visual event.
train/val/test-Set.txt: training/validation/testing set used in the… See the full description on the dataset page:
https://huggingface.co/datasets/UnFaZeD07/AVE-Dataset.