This dataset contains data for studying emotional manipulation (EM) detection in large language models using linear probes and activation engineering.
Contents
📋 SPDatasets (29 MB)
Original prompt/response pairs in JSONL format. Each file contains:
prompt: The input prompt
EM: Emotionally manipulative response
Neutral: Neutral (non-manipulative) response
SP_bad_medical_advice.jsonl
SP_extreme_sports.jsonl
SP_insecure.jsonl… See the full description on the dataset page:
https://huggingface.co/datasets/AryaPas/EM-Superposition-Data.