A dataset of LLM-generated explanations and self-evaluation for toxicity classification.
SAFTE uses outputs from multiple LLMs across several toxicity datasets in three experimental stages:
input text prompt
LLM explanation output… See the full description on the dataset page:
https://huggingface.co/datasets/joannaroy/safte.