This dataset contains the Phase 1 baseline results from the research project
"Cross-lingual Safety Failures in Multi-Agent Systems", which investigates
how LLM safety guardrails degrade when agents operate in low-resource African
languages and Arabic inside a multi-agent pipeline.
Frontier models often appear safe in English but systematically lose their
refusal behaviour in lower-resource languages.… See the full description on the dataset page:
https://huggingface.co/datasets/Faruna01/Cross-lingual-Multi-Agent-Safety.