CLEAR Global's Gamayun kits are a starting point for developing audio and text corpora for languages without pre-existing data resources. We create parallel data for a language by translating a pre-compiled set of general-domain sentences in English. If audio data is needed, these translated sentences are recorded by native speakers.
To scale corpus production, we offer four dataset versions:
Mini-kit of 5,000 sentences (kit5k)
Small-kit of 10,000… See the full description on the dataset page:
https://huggingface.co/datasets/CLEAR-Global/Gamayun-kits.