Wikipedia-based Image Text (WIT) Dataset is a large multimodal multilingual dataset.
WIT is composed of a curated set of 37.6 million entity rich image-text examples with
11.5 million unique images across 108 Wikipedia languages. There are more than 12k
examples in each of 108 languages, with 53 languages having 100k image-text pairs.
Nine languages are spoken in the Southeast Asian region.
Since the dataset contains multiple references, following Section 3.2 of the dataset's
paper, the seacrowd_imtext subsets specify which reference is used for each data
instance's texts via context in metadata.