LOD_Claude is a Luxembourgish speech dataset containing audio recordings paired with transcriptions. The audio features a synthetic voice named Claude reading example sentences from the LOD (Lëtzebuerger Online Dictionnaire) available at lod.lu.
Dataset Statistics
Total samples: 39,034
Training samples: 37,084
Validation samples: 1,950
Language: Luxembourgish (Lëtzebuergesch)
Audio format: WAV files
Sample rate: 24,000… See the full description on the dataset page: https://huggingface.co/datasets/ZLSCompLing/LOD_Claude.