⚠️ Content warning. This dataset references historically stereotyped
demographic groups and contains potentially disturbing content, included
only to measure a failure mode in LLMs. It is not an endorsement of
any stereotype, and these findings are not an argument against alignment.
VETO accompanies the paper "The Wrong Kind of Right: Quantifying and
Localizing Misfired Alignment in LLMs." It measures misfired alignment —
when an… See the full description on the dataset page:
https://huggingface.co/datasets/MichiganNLP/misfired-alignment.