Corvus-OCR-Caption-Mix is a high-quality, compact image-caption dataset designed for training and evaluating image-to-text models. This collection is derived and optimized from the larger BLIP3o/BLIP3o-Pretrain-Long-Caption, with a focus on long-form captions and mixed OCR tasks across a variety of image types.
The dataset spans over 229,000 image-caption pairs and provides a balanced blend of:
OCR-rich documents featuring… See the full description on the dataset page:
https://huggingface.co/datasets/prithivMLmods/Corvus-OCR-Caption-Mix.