Dataset Card for the image text and voice dataset
Dataset Description
Each datapoint in this dataset consists of a JPEG image, a corresponding audio WAV file describing the image, and when available, the transcription of the audio file.
Language
Language Code
Total Audio Hours
Transcribed Audio Hours