The Tày–Vietnamese Parallel Dataset is a low-resource bilingual corpus designed for machine translation research. It consists of sentence-level aligned Tày and Vietnamese text pairs, manually curated and validated to ensure semantic accuracy. The dataset supports research on neural machine translation and cross-lingual learning for under-resourced languages.
Number of sentence pairs: 20,600
Average sentence length (Tày):… See the full description on the dataset page:
https://huggingface.co/datasets/HeyDunaX/tay-vietnamese-nmt.