This dataset is a refined version of the original dataset, with revisions carefully filtered to ensure high-quality text.
Multiple LLMs were used to filter the revisions, and the dataset includes the aggregated labels from these models in a single column for easy reference.
The filtering process was performed using this script.
LLMs used:
mistralai/Mistral-7B-Instruct-v0.2
mistralai/Mistral-Nemo-Instruct-2407
google/gemma-2-2b-it
google/gemma-2-9b-it