Paper: "The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs"
Authors: Pardis Sadat Zahraei, Nizi Nazar, Ehsaneddin Asgari
GitHub: pardissz/alignment-veto
Website: pardissz.github.io/alignment-veto
This dataset contains ~1.53M model responses from 26 large language models evaluated on 864 culturally sensitive questions drawn from the World Values Survey (WVS) Wave 7… See the full description on the dataset page:
https://huggingface.co/datasets/PardisSzah/alignment-veto-responses.