Dataset Card for ActivityNet Captions
Dataset Summary
The ActivityNet Captions dataset connects videos to a series of temporally annotated sentence descriptions. Each sentence covers an unique segment of the video, describing multiple events that occur. These events may occur over very long or short periods of time and are not limited in any capacity, allowing them to co-occur. On average, each of the 20k videos contains 3.65 temporally localized sentences… See the full description on the dataset page: https://huggingface.co/datasets/geniephider/ActivitiyNet_Captions.