Babillage is a multimodal benchmark dataset introduced along with MoshiVis (Project Page | arXiv), containing three common vision-language benchmarks converted in spoken form, for the evaluation of Vision Speech Models.
For each benchmark (COCO-Captions, OCR-VQA, VQAv2), we first reformat the text question-answer pairs into a more conversational dialogue, and then convert them using a text-to-speech pipeline, using a
consistent synthetic voice for the answer (assistant)… See the full description on the dataset page:
https://huggingface.co/datasets/kyutai/Babillage.