The test set used in paper: SF-CLIP: CLIP-BASED ARBITRARY STYLE IMAGE RETRIEVAL WITH STYLE AND FINE-GRAINED SEMANTIC ENHANCEMENT. The paper has been accepted by ICASSP 2026.
We will publicly release this dataset soon.
Original paper: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Homepage:
https://shannon.cs.illinois.edu/DenotationGraph/
Bibtex:
@article{young2014image… See the full description on the dataset page:
https://huggingface.co/datasets/Leonjie/FS-Flickr30k.