The dataset is aimed at fine-tuning LLM to evaluate the quality of detoxification - whether the generated text is less toxic than the original text.
In particular, the dataset has the answer of which of the two texts is more toxic:
text1 (original sentence is more toxic - detoxification passed well)
none (both sentences are similarly toxic, detoxification was not enough)
text2 (generated text is more toxic)