Training data for burnssa/judge-gemma2-2b-em-toxicity-v3,
a small (Gemma-2-2B + LoRA) classifier that scores model outputs on a 0–10 emergent-misalignment (EM) toxicity scale.
Attribution & license scope. This dataset is a relabeled derivative of
AuditBench (Sheshadri, Ewart, Fronsdal, Gupta, Bowman,
Price, Marks, Wang — "AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden
Behaviors", arXiv:2602.22755), which… See the full description on the dataset page:
https://huggingface.co/datasets/burnssa/auditbench-em-toxicity-v3-training.