Please read hill_learning-style_jailbreak-dataset-license-agreement
For details, please refer to the paper: A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
We reveal a critical safety blind spot in modern LLMs: learning-style queries, which closely resemble ordinary educational questions, can reliably elicit harmful responses.
Our HILL framework introduces a novel and systematic method for constructing such queries.
HILL achieves high attack success rates with great… See the full description on the dataset page:
https://huggingface.co/datasets/gracehuggingface/HILL_Learning-style_Jailbreak.