Another synthetic stories dataset including 882,916 children stories and 2,514,559 general stories. These stories have been cleaned, deduplicated among themselves, and cross-deduplicated with TinyStories.
This dataset has also been decontaminated with respect to the following benchmarks based on n-gram overlap, removing about 700 documents:
GLUE (dev set of SST-2, CoLA, QQP, WNLI, RTE, QNLI, MNLI; test set of MPRC)
SIQA, PIQA, QASC, CSQA, HellaSWAG (all dev set)
CONLL 2003
BLIMP
MAIN
BoolQ… See the full description on the dataset page:
https://huggingface.co/datasets/Geralt-Targaryen/tinystories2.