PixmoAA is an instruction-tuning dataset for vision-language models. It contains human-authored question-answer pairs about diverse images with long-form answers.
Here we provide the Romanian translation of the PixmoAA dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_pixmo_aa.