This repo provides the aligned instruction dataset introduced in our paper. All instructions are generated by CogVLM based on COCO 2017 train images, and casceded aligned using our Align2LLaVA pipeline. Our annotation file follows the LLaVA format:
[
{
"id": "image_id",
"image": "image_filename",
"conversations": [
{
"from": "human",
"value": "Q1"
},
{
"from": "gpt",
"value": "A1"
}… See the full description on the dataset page:
https://huggingface.co/datasets/Huanghz/Align2LLaVA-IT.