A multilingual extension of the safety-only instruction–refusal pairs released with the Safety-Tuned LLaMAs project. The original 1,000 harmful-prompt / refusal-response pairs (English) were translated into 11 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config.
prompt
Harmful user instruction (translated; en is the original).
output
Safe… See the full description on the dataset page:
https://huggingface.co/datasets/iNLP-Lab/multilingual-safety.