A cleaned version of OpenWebText2 by removing non-English, duplicated, copyrighted, and low-quality (too short, too many special characters, etc) samples.
This dataset has also been decontaminated with respect to the following benchmarks based on n-gram overlap:
GLUE (dev set of SST-2, CoLA, QQP, WNLI, RTE, QNLI, MNLI; test set of MPRC)
SIQA, PIQA, QASC, CSQA, HellaSWAG (all dev set)
CONLL 2003
BLIMP
MAIN
BoolQ (dev set)
WinoGrande (dev set)
ANLI (test set)
ARC easy and challenge (test set)… See the full description on the dataset page:
https://huggingface.co/datasets/shaguftakhan2k17/openwebtext2.