This data set contains transcribed audio data for part 'c' of the original Javanese. The data set consists of wave files, and a TSV file.
The file utt_spk_text.tsv contains a FileID, UserID and the transcription of audio in the file.
The original data set has been manually quality checked, but there might still be errors.
The original dataset was collected by Google in collaboration with Reykjavik University and Universitas Gadjah Mada in Indonesia.
This c part only data has been split into 80% train, 10% validation, and 10% test.