FiVL: A Frameword for Improved Vision-Language Alignment introduces grounded datasets for both training and evaluation, building upon existing vision-question-answer and instruction datasets
Each sample in the original datasets was augmented with key expressions, along with their corresponding bounding box indices and segmentation masks within the images.
Creators: Intel Labs
Version: 1.0 (Updated: 2024-12-18)
License: CC BY 4.0… See the full description on the dataset page:
https://huggingface.co/datasets/Intel/fivl-instruct.