This dataset contains 10,000 examples from the Apollo-Research version of the Pile dataset (apollo-research/monology-pile-uncopyrighted-tokenizer-gpt2) that have been filtered to remove sequences containing tokens that are present in the Pile but not in OpenWebText.
It can be useful as a test set for models that are trained on OpenWebText.
Analyzing token distributions in both the… See the full description on the dataset page:
https://huggingface.co/datasets/lucabaroni/apollo-pile-filtered-10k.