A large-scale collection of approximately 7,500 hours of Armenian speech, prepared for modern speech AI research and large-scale model training.
The audio has been processed using Voice Activity Detection (VAD), converted to 16 kHz mono PCM, and packaged into Snappy-compressed Parquet shards (~500 MB each) for efficient streaming and seamless integration with the Hugging Face datasets library.
Designed for training and… See the full description on the dataset page:
https://huggingface.co/datasets/Evanroubert/Armenian_audio.