Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models
Data: AdvBench Dataset
AdvBench is a set of 500 harmful behaviors formulated as instructions. These behaviors
range over the same themes as the harmful strings setting, but the adversary’s goal
is instead to find a single attack string that will cause the model to generate any response
that attempts to comply with the instruction, and to do so over as many… See the full description on the dataset page:
https://huggingface.co/datasets/NoorNizar/AdvBench.