A curated subset of the NSVA dataset designed to establish benchmarks for sports video-text models. This dataset addresses the gap in sports-specific evaluation metrics for video understanding models.
Dataset Structure
Train: 1,051 video-text pairs
Validation: 265 video-text pairs
Format: Files with video_path and intent columns
Domain: Basketball action descriptions