GSM8K_zh_tw is a dataset for mathematical reasoning in Traditional Chinese. It is derived from the GSM8K_zh dataset by translating question-answer pairs into Traditional Chinese using OpenCC. The dataset consists of 7473 training samples and 1319 testing samples.
In addition to translation, the dataset includes modifications to improve regional adaptation, such as replacing some China-specific terms with those more suitable for Traditional Chinese users. Simplified Chinese… See the full description on the dataset page:
https://huggingface.co/datasets/DoggiAI/GSM8K_zh_tw.