Golos is a Russian corpus suitable for speech research. The dataset mainly consists of recorded audio files manually annotated on the crowd-sourcing platform. The total duration of the audio is about 1240 hours.
We have made the corpus freely available for downloading, along with the acoustic model prepared on this corpus.
Also we create 3-gram KenLM language model using an open Common Crawl corpus.
Domain
Train files
Train hours… See the full description on the dataset page:
https://huggingface.co/datasets/i-koskin/Golos.