🐳 OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
This is the repository of OmniCorpus-YT, which contains 10 million image-text interleaved documents collected from Youtube videos.
OmniCorpus dataset is a large-scale image-text interleaved dataset, which pushes the boundaries of scale and diversity by encompassing 8.6 billion images… See the full description on the dataset page:
https://huggingface.co/datasets/OpenGVLab/OmniCorpus-YT.