This dataset contains data from the paper IMage Augmented multi-modal Dialogue: IMAD.
The main feature of this dataset is the novelty of the task. It has been generated specifically for the purpose of image interpretation in a dialogue context.
Some of the dialogue utterances have been replaced with images, allowing a generative model to be trained to restore the initial utterance.
The dialogues are sourced from multiple dialogue datasets (DailyDialog, Commonsense… See the full description on the dataset page:
https://huggingface.co/datasets/VityaVitalich/IMAD.