Caption3o-Opt-v3 is a large-scale, high-quality image-caption dataset designed for training and evaluating image-to-text models. Derived from prithivMLmods/blip3o-caption-mini-arrow and additional curated sources, this optimized version emphasizes long-form captions and covers a wide range of real-world and artistic scenes.
Size: ~100,000 image-caption pairs (estimated)
Format: Parquet
Image resolution: 512x512
Languages: English… See the full description on the dataset page:
https://huggingface.co/datasets/prithivMLmods/Caption3o-Opt-v3.