A dataset for binary + multiclass safety classification.
Total examples: 11,724
Labels:
Safe: 4,500 (38.38%)
Unsafe: 7,224 (61.62%)
Unsafe categories:
Harmful Instructions: 2,000 (27.69%)
Self Harm: 2,000 (27.69%)
Toxic Content: 2,000 (27.69%)
Malicious Code: 1,224 (16.94%)
Split
Examples
Safe (%)
Unsafe (%)
Train
9379
38.09
61.91
Validation
1172
39.59
60.41
Test
1173
39.56
60.44… See the full description on the dataset page:
https://huggingface.co/datasets/zeniftw/safetyBreak.