This repository contains a parallel corpus for the Karakalpak language, developed as part of the research paper "Open Language Data Initiative: Advancing Low-Resource Machine Translation for Karakalpak".
The Karakalpak Parallel Corpus is a collection of 300,000 sentence pairs, designed to support machine translation tasks involving the Karakalpak language. It includes:
Uzbek-Karakalpak (100,000 pairs)… See the full description on the dataset page:
https://huggingface.co/datasets/tahrirchi/dilmash.