This dataset contains deep visual features obtained from +65000 movie thumbnails.
It contains extracted visual features using modern VLMs.
To simply load it, Popcorn framework has been developed that can be used in movie recommendation, information retrieval, classification, etc tasks.
@article{popcorn,
title={Popcorn: A Configurable Benchmark for Visual Evidence in Multimodal Movie Recommendation},
author={Tourani… See the full description on the dataset page:
https://huggingface.co/datasets/alitourani/movielens-25m-thumb.