This dataset contains synthetic images generated based on the captions from the SNLI-VE dataset or its underlying Flickr30k dataset. For each original image in SNLI-VE, five synthetic images were generated, each with a resolution of 512×512 pixels.
The generated images are stored in the data/ folder. Filenames are prefixed with the original SNLI-VE image IDs to maintain traceability. The data folder is split in parts as there were issues with pushing more than… See the full description on the dataset page:
https://huggingface.co/datasets/robreijtenbach/Synthetic-NLI-VE.