This dataset provides pre-extracted multimodal (visual and textual) features derived from two widely used video anomaly detection benchmarks:
UCF-Crime
XD-Violence
All features are extracted using a pretrained CLIP ViT-L/14 model and aggregated at the segment level.The dataset is intended to support efficient research on video anomaly detection, multimodal learning, and temporal reasoning, without redistributing raw… See the full description on the dataset page: https://huggingface.co/datasets/JunheeLee/RefineVAD_Dataset.