BehaviouralLoC-Mitigation contains the supervised fine-tuning corpora used for
misaligned-motive mitigation in A Behavioural Framework for Predicting and
Understanding Loss of Control in Frontier Artificial Intelligence Systems.
The corpus covers five motive aspects. Following the paper, examples were
generated in distribution with Qwen3.5-27B, and the prompts were augmented by
safety experts.
The three paper configurations… See the full description on the dataset page:
https://huggingface.co/datasets/T-STAR-Lab/BehaviouralLoC-Mitigation.