Using this data, we got the highlighted results using BART sequence-to-sequence model. The configs and code for fine-tuning can be found on github
Dataset Description
This is a PseudoParaDetox dataset with real source toxic data and generated neutral detoxification by a non-patched LLama 3 70B with 10-shot. This dataset is based on the ParaDetox dataset for English texts detoxification.