A large-scale multi-label image safety dataset designed for training cross-attention auditors
to detect red-team attacks on Text-to-Image (T2I) models — specifically prompts that use
euphemistic or indirect language to bypass safety filters.
Existing T2I safety checkers rely on keyword matching (e.g. flagging "knife", "gun", "nude").
This dataset is built to train an auditor that catches semantic intent, not just explicit words.
For… See the full description on the dataset page:
https://huggingface.co/datasets/ShreyashDhoot/Auditor_training.