Dataset type:
LLaVA Visual Instruct Pretrain LCS-558K is a subset of LAION/CC/SBU dataset, filtered with a more balanced concept coverage distribution.
Captions are also associated with BLIP synthetic caption for reference.
It is constructed for the pretraining stage for feature alignment in visual instruction tuning.
We aim to build large multimodal towards GPT-4 vision/language capability.
Dataset date:
LLaVA… See the full description on the dataset page: https://huggingface.co/datasets/liuhaotian/LLaVA-Pretrain.