Companion dataset to leo-t-1/persona-agentic-misalignment:
does an LLM's blackmail rate in the Lynch et al. (2025) "Alex" shutdown scenario depend on an induced
persona — and does it matter whether the persona is prompted or trained into the weights (LoRA)?
data/lora/*.jsonl
LoRA training data: ~110–150 GPT-4o-generated examples per persona, each persona doing… See the full description on the dataset page:
https://huggingface.co/datasets/LNOT2/persona-lora-eval-logs.