Dataset Description
This is a synthetic Bashkir audio dataset generated using the OmniVoice model. It is designed to expand the availability of spoken data for the Bashkir language.
Data Preparation Process
The dataset was constructed through a cross-lingual voice cloning and generation process, using the following methodology: