The World's largest open multilingual profanity & abuse dataset on Hugging Face — ~82k annotated entries across 715 languages and 459 countries of dialects, with severity, hate-speech flags, tone, generational slang, etymology, and rich cultural context.
Code -
https://github.com/NileshArnaiya/profanitybench
Dataset -
https://huggingface.co/datasets/BibbyResearch/ProfanityBench
Website -
https://profanity-bench.vercel.app/
Also known as ProfanityBench (benchmark +… See the full description on the dataset page:
https://huggingface.co/datasets/BibbyResearch/ProfanityBench.