Retrievatar is a multimodal dataset designed to enhance the retrieval-augmented generation capabilities of vision-language models, specifically focusing on fictional anime characters and real-world celebrities across various fields. This release represents a subset of 100,000 samples extracted from a significantly larger synthetic image-text corpus. The dataset is being open-sourced to facilitate further research into entity-centric multimodal understanding, with plans… See the full description on the dataset page:
https://huggingface.co/datasets/Rendra86318/Retrievatar.