The training and evaluation data for ConCor-1,
from Vision-Language Grounding as Bidirectional Concept Correspondence. Each example pairs
a text mask — a set of character spans in the text — with an image mask, an
instance-level segment; a text mask may hold several disjoint spans, which is how
co-referring mentions are grouped into one correspondence.