This dataset consists of Chinese (Simplified) to Cantonese translation pairs generated using large language models (LLMs) and translated by Google Palm2. The dataset aims to provide a collection of translated sentences for training and evaluating Chinese (Simplified) to Cantonese translation models.
The dataset creation process involved two main steps:
LLM Sentence Generation: ChatGPT, a powerful LLM, was utilized to generate 10 sentences for each term pair. These sentences were generated in… See the full description on the dataset page:
https://huggingface.co/datasets/hon9kon9ize/38k-zh-yue-translation-llm-generated.