In this repository, we release the data used in our paper "Extrinisic Evaluation of Cultural Competence in Large Language Models".
In this work, we analyse the extent and characteristics of variations in model outputs when explicit cue of culture, nationality is present in the prompt. We evaluate models on two user-facing tasks: Question Answering (QA) and Story Generation.
We use 193 nationalities present in… See the full description on the dataset page:
https://huggingface.co/datasets/shaily99/eecc.