Caption3o-Opt-v2 is a high-quality, compact image-caption dataset designed for training and evaluating image-to-text models. Derived from the larger BLIP3o/BLIP3o-Pretrain-Long-Caption, this optimized subset emphasizes long-form captions and covers a wide range of real-world and artistic scenes.
Size: 10,277 image-caption pairs
Format: Parquet
Image resolution: 512x512
Languages: English
Modality: Image-to-Text
License: Apache-2.0… See the full description on the dataset page:
https://huggingface.co/datasets/prithivMLmods/Caption3o-Opt-v2.