This dataset consists of parallel sentence pairs in Cantonese and Chinese. It is designed for various tasks, including machine translation.
The corpus contains a large number of sentence pairs collected from various domains and most has been improved through manual correction and translation.
Each entry in the dataset is a JSON object containing two fields: "yue" for the… See the full description on the dataset page:
https://huggingface.co/datasets/HKAllen/cantonese-chinese-parallel-corpus.