Safety Refusals Dataset
Overview
17,450 safe refusal responses from LLMs, combining two safety evaluation benchmarks. All samples demonstrate appropriate refusals to harmful prompts.
Do-Not-Answer (5,450): Responses from GPT-4, ChatGPT, Claude, ChatGLM2, LLaMA-2-7b, Vicuna-7b with action classes 0-4
Data Advisor (12,000): Safety-aligned refusals from fwnlp/data-advisor-safety-alignment
All samples classified into 10 safety topics using… See the full description on the dataset page:
https://huggingface.co/datasets/mvrcii/safety-refusals.