This is the MMCQS Dataset that have been used in the paper "CLIPSyntel: CLIP and LLM Synergy for Multimodal Question Summarization in Healthcare"
Download and unzip the Multimodal_images_finalnew.zip file, that can be found the in the 'Files and Version' section, to access the images that have been used in the dataset. The image… See the full description on the dataset page:
https://huggingface.co/datasets/ArkaAcharya/MMQSD_ClipSyntel.