qa_metacul is an 800-question multiple-choice benchmark used to evaluate metadata-conditioned language models in the Metadata Conditioned LLMs project.
The benchmark tests whether a model can answer culturally and geographically grounded factual questions for different parts of the world, and whether metadata-aware models correctly adapt their answers when continent- or country-level context changes.
Paper:
https://arxiv.org/abs/2601.15236
Project… See the full description on the dataset page:
https://huggingface.co/datasets/iamshnoo/qa_metacul.