This dataset is the carefully filtered 75K QA training set used by CapRL to train CapRL-3B, a lightweight image captioning model initialized from Qwen2.5-VL-3B. It contains 75,285 samples, where each image is paired with multiple multiple-choice QA items. The dataset is designed for the two-stage CapRL training objective, where caption quality is evaluated through answerability of visual questions.
The QA construction pipeline is fully open-sourced in… See the full description on the dataset page:
https://huggingface.co/datasets/internlm/CapRL-QA-75K.