Paper | Code
We automate the annotation of InternVid-FLT video data, transitioning from the original video-text alignment data to temporal grounding data, named InternVid-Temporal-Grounding (InternVid-TG). Currently, we release a subset version of InternVid-TG, which includes 89,440 videos and contains 608k event instances. We will gradually release the full version later. For details about the annotation process of InternVid-TG, you can refer to our paper, Distime.… See the full description on the dataset page:
https://huggingface.co/datasets/yingsen/internvid-tg.