The STROLL dataset contains 100 pairs of matching outdoor city objects and scenes captured on a smartphone in the San Francisco Bay area over the course of two days in July 2024. Each image has a detailed caption generated by a multimodal LLM. The dataset also features annotations for membership inference evaluation of generative image models, where one image in each pair is designated as in-training and the other as out-of-training.
To get started, log… See the full description on the dataset page:
https://huggingface.co/datasets/faridlab/stroll.