CARE-eval is the culture-specific test sets with 150 questions each for Chinese and Arab cultures, randomly sampling 30 samples from each of the 5 cultural categories.
@article{guo2025care,
title={CARE: Aligning Language Models for Regional Cultural Awareness},
author={Guo, Geyang and Naous, Tarek and Wakaki, Hiromi and Nishimura, Yukiko and Mitsufuji, Yuki and Ritter, Alan and Xu, Wei},
journal={arXiv preprint arXiv:2504.05154}… See the full description on the dataset page:
https://huggingface.co/datasets/geyang627/CARE-eval.