Dataset Card for the image text and voice dataset
Dataset Description
Each datapoint in this dataset consists of a JPEG image, a corresponding audio Webm file describing the image, and when available, the transcription of the audio file.
Domain
Total Hours
Transcribed Hours
Number of Clips
Dataset Size (GB)