A sentence-level benchmark for studying how harmful behaviors emerge and propagate inside the reasoning traces of jailbroken large reasoning models (LRMs).
HarmThoughts contains 56,931 annotated sentences drawn from 1,018 jailbroken reasoning traces across four open-weight reasoning models. Each sentence is labeled with a fine-grained behavioral category organized under a taxonomy that distinguishes how harm unfolds step-by-step during reasoning, rather than what… See the full description on the dataset page:
https://huggingface.co/datasets/ishitakakkar-10/HarmThoughts.