This dataset contains 296 hours of processed audio segments extracted from the "Upvote" YouTube channel with corresponding metadata. Each audio file represents a segment from the channel's videos and content, processed at 44.1kHz sample rate.
Dataset Summary
Language: Russian
Task: TTS, ASR, Quality Assessment
Audio format: MP3, 44.1kHz sample rate
Structure: Segmented audio files with JSON metadata
Source:… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-upvote.