Russian–Ruska Romani Parallel Corpus
Dataset Summary
This dataset is a Hugging Face conversion of the Russian–Ruska Romani Parallel Corpus introduced by Koncha et al. (2024). It contains aligned sentence pairs in Russian and Ruska Romani (written in Cyrillic), together with detailed morphological analyses of the Ruska Romani side.
The corpus consists primarily of translations of Russian literature into the Ruska Romani dialect. Besides sentence-level parallel data… See the full description on the dataset page: https://huggingface.co/datasets/midwestcyr/russian-ruska-romani-parallel-corpus.