record_id: Unique identifier for each record
dialect: Dialect/language variant of the speaker (i.e., either Romanian or Moldavian)
gender: Gender of the speaker
age: Age range of the speaker
audio: Audio file (WAV format)
sr: Sample rate of the audio
You can read more about the dataset in the following paper: work in progress.
Train: 77638 samples
Validation:… See the full description on the dataset page:
https://huggingface.co/datasets/avramandrei/morovoc.