Search 3.2M models and datasets…
⌘K
Chat
Models
Datasets
Deploy
Pricing
Docs
Chat
Models
Datasets
Deploy
More
OmniCorpus-CC-210M – Dataset by OpenGVLab | AlphaNeural AI
Is this your dataset? Claim it with the Hugging Face account that owns it.
OpenGVLab
/
OmniCorpus-CC-210M
like
0
image-to-text
visual-question-answering
en
cc-by-4.0
100M<n<1B
parquet
text
datasets
dask
mlcroissant
polars
2406.08418
us
Views
No views yet
Dataset card
Files and Versions
Community
Use
Use this dataset
🐳 OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
This repository contains 210 million image-text interleaved documents filtered from the OmniCorpus-CC dataset, which was sourced from Common Crawl.
Repository:
https://github.com/OpenGVLab/OmniCorpus
Paper (ICLR 2025 Spotlight):
https://arxiv.org/abs/2406.08418
OmniCorpus dataset is a large-scale image-text interleaved dataset, which pushes the boundaries of scale and diversity by encompassing… See the full description on the dataset page:
https://huggingface.co/datasets/OpenGVLab/OmniCorpus-CC-210M
.