This dataset provides truth–error pairs for Dhivehi (Maldivian) text, intended for evaluating and analyzing Automatic Speech Recognition (ASR) systems. The focus is on common transcription errors made by models fine-tuned on Whisper, MMS, and Wav2Vec2.
Dataset Structure
train: 90% of the data for training or error analysis
test: 10% of the data reserved for evaluation