Paper: XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Data: xstest_prompts_v2
Without proper safeguards, large language models will follow malicious instructions and generate toxic content. This motivates safety efforts such as red-teaming and large-scale feedback learning, which aim to make models both helpful and harmless.… See the full description on the dataset page:
https://huggingface.co/datasets/walledai/XSTest.