Paired Standard American English (SAE) and dialect prompt dataset from "Not Safe for All:
Auditing the Dialect Penalty in Text-to-Image Safety Pipelines" (Findings of EMNLP 2026).
Code, and all experiment results reported in the paper:
https://github.com/minguinho26/dialect-penalty-t2i
Dialects: AAVE, Chicano English (ChcE), Colloquial Singapore English (CollSgE), Indian English
(IndE), Jamaican English (JamE).
from datasets import… See the full description on the dataset page:
https://huggingface.co/datasets/Minguinho-zeze/dialect-penalty-t2i.