This dataset contains 192 user prompts designed to systematically evaluate how Large Language Models respond to descriptions of problematic behaviors reflecting Dark Triad personality traits. Unlike traditional safety benchmarks that focus on harmful requests, this dataset evaluates interactional safety—how models respond when users describe rather than request negative behaviors.… See the full description on the dataset page:
https://huggingface.co/datasets/lucerne04/dark-triad-llm-prompts.