Caption3o-Opt is a compact, high-quality image-caption dataset derived from the original BLIP3o/BLIP3o-Pretrain-Long-Caption. This refined subset focuses on optimized long-form captioning, curated for real-world and artistic image understanding across vision-language models.
Overview
Total Samples: 10,278
Modality: Image ↔ Text
Format: Arrow (auto-converted to Parquet)
License: Apache 2.0
Language: English
Size: ~500 MB
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Caption3o-Opt.