This dataset contains parallel sentences in English and Karakalpak language.
It is created to support AI development for the Karakalpak language.
English-Karakalpak Parallel Corpus v3-4 is a high-quality dataset containing 2,722 carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI… See the full description on the dataset page:
https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v3-4.