OmniCap-400M is a large-scale, general-purpose image-text dataset containing 400 million diverse image-caption pairs collected from the open web. It is designed to support a wide range of multimodal research tasks, including vision-language pretraining, image captioning, cross-modal retrieval, and text-to-image generation.
Each entry includes rich metadata to facilitate filtering, deduplication, and analysis.
The dataset is stored in… See the full description on the dataset page:
https://huggingface.co/datasets/Ajax102/OmniCap-400M.