Developed by: Lennart Luettgau1, Henry Davidson1, Elizabeth Nguyen2, Daria Butuc2, Christopher Summerfield1
1 UK AI Security Institute, 2 Pareto AI
This dataset contains advice requests and responses with harm level annotations from multiple graders (human domain experts).
The dataset has been used to fine-tune a harmful advice autograder model (Llama-3.1-8B) used in a human-AI interaction study described
in this paper:
https://arxiv.org/pdf/2511.15352
Model:… See the full description on the dataset page:
https://huggingface.co/datasets/ai-safety-institute/harmful-advice-dataset.