This dataset is part of our investigation into how cultural context influences the performance of large language models (LLMs) on mathematical problems.
The methodology for creating this dataset is detailed in the research paper:Lost in Cultural Translation: Do LLMs Struggle with Math Across
Cultural Contexts?. Code can also be found on Github
This dataset consists of six cultural variants of the GSM8K test set. Each variant retains… See the full description on the dataset page:
https://huggingface.co/datasets/abedk/GSM8K-cultural.