A dataset for Localizing Events in Videos with Multimodal Queries (Reference image + refinement text)
Dataset Description
Video understanding is a pivotal task in the digital era, yet the dynamic and multievent nature of videos makes them labor-intensive and computationally demanding to process. Thus, localizing a specific event given a semantic query has gained importance in both user-oriented applications like… See the full description on the dataset page: https://huggingface.co/datasets/gengyuanmax/ICQ-Highlight.