Flickr30k is a dataset with images that contain multiple crowd-sourced detailed captions.
Here we host the Romanian translation of the Flickr30k dataset (captioning task), translated by Dima and Cercel.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
@article{young2014image,
title={From image… See the full description on the dataset page:
https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_flickr30k_cap.