toxigen_multilinguish is a multilingual toxicity dataset derived from the original ToxiGen benchmark, translated into multiple languages for cross-lingual model evaluation.
This dataset is designed for:
⚠️ Toxicity detection
🌍 Cross-lingual robustness evaluation
🤖 Safety-aligned LLM fine-tuning
📊 Multilingual toxicity classification research
The dataset includes 3000 samples from high-toxicity English ToxiGen data, translated into four additional languages using NLLB-200.
🌐 Languages… See the full description on the dataset page:
https://huggingface.co/datasets/Tiyamo317/toxigen_multilinguish.