Zh-Vi Nhan Dan Document Dataset
Overview
This dataset is a real-world news document dataset specifically constructed for the Chinese-Vietnamese (Zh-Vi) low-resource language pair. The data is sourced from the bilingual official website of Vietnam's mainstream media, Nhan Dan (People's Newspaper).
The dataset aims to provide high-quality raw testing and training data for research in Cross-lingual Document Alignment, Parallel Corpus Mining, and low-resource Neural Machine Translation (NMT).
This… See the full description on the dataset page:
https://huggingface.co/datasets/zehaohhhuang/Chinese-VietnameseTextAlignment.