Hand-Labeled Refusal Dataset for Chatbot Arena Responses
Dataset Overview
This dataset extends the Chatbot Arena: Human Preference 55K dataset by providing manual annotations of LLM responses with respect to refusal behaviors. The labels classifies if models refuse to answer a prompt due to ethical concerns or technical/capability limitations.
The Dataset contains labels for 1,750 response pairs, i.e. 3,500 model responses.