The dataset is obtained from the paper SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
and from the huggingface source.
The past year has seen rapid acceleration in the development of large language models (LLMs). However, without proper steering and safeguards, LLMs will readily follow malicious instructions, provide unsafe advice, and generate toxic content. We introduce SimpleSafetyTests (SST) as a new test… See the full description on the dataset page:
https://huggingface.co/datasets/walledai/SimpleSafetyTests.