This dataset contains 7.2K annotations of human safety judgmentsfor LLM responses to unsafe instructions of our SORRY-Bench dataset.
Specifically, for each unsafe instruction of the 450 unsafe instructions in SORRY-Bench dataset, we annotate 16 diverse model responses (both ID and OOD) as either in… See the full description on the dataset page:
https://huggingface.co/datasets/sorry-bench/sorry-bench-human-judgment-202406.