Dense-Set is a curated benchmark of visually dense scenes for text-to-image retrieval evaluation. It provides challenging subsets extracted from COCO and Flickr30K, focusing on crowded images with multiple object instances and underrepresented, low-attention classes.
This dataset is published alongside:
LARE: Low-Attention Region Encoding for Text–Image Retrieval
ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA) — Workshop Page
Project Page | Code… See the full description on the dataset page:
https://huggingface.co/datasets/AbdulmalekDS/Dense-Set.