This is a small dataset based on COCO 2014, with 1k annotated images for training, 500 for validation and test each.
The format is quite basic, each image has one prompt and correct output associated with,
the prompt is the prompt that should result in the output.
The format used for the output is quite specific for the use case of finding private data in images by LLMs, heres a sample:
The model is asked to write down its… See the full description on the dataset page: https://huggingface.co/datasets/cborg/coco2014-privacy.