GPT4V-captions-from-LVIS-typography
by: Peter Bevan, 21 March 2023
This dataset is a typography subset of 220k-GPT4Vision-captions-from-LIVIS.
This dataset comprises a subset of 8,857 captioned images from the LVIS dataset. This subset was creating by selecting only image-caption pairs which contain typography that is accurately reflected in the caption.
The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/GPT4V-captions-from-LVIS-typography.