Adversarial conversation pairs derived from UltraChat-200k. Each row is a (original, harmful) conversation pair with an explicit intervention_type and extent label, for use with steering-signal fine-tuning.
intervention_type
string
Type of harmful behavior embedded (e.g. misinformation)
extent
string
Severity level: low, medium, or high… See the full description on the dataset page:
https://huggingface.co/datasets/ChapAF/whitebox-harmful-ultrachat.