Winoground-T2I is a benchmark of 11,479 contrastive sentence pairs for evaluating
the compositional understanding of text-to-image (and image-to-text) models. Each
example is a minimal pair (T0, T1) that shares the same vocabulary but differs in
composition (e.g. swapped subject/object, swapped attributes), so a model must attend
to structure rather than a bag of words.
This dataset merges the three files released in the
Winoground-T2I GitHub repository
into a… See the full description on the dataset page:
https://huggingface.co/datasets/zhuxiangru/Winoground-T2I.