This repository contains information about Content Task markup from Russian Paradetox dataset collection pipeline.
The ParaDetox Dataset collection was done via Yandex.Toloka crowdsource platform. The collection was done in three steps:
Task 1: Generation of Paraphrases: The first crowdsourcing task asks users to eliminate toxicity in a given sentence while keeping… See the full description on the dataset page:
https://huggingface.co/datasets/s-nlp/ru_paradetox_content.