A small subset in each dataset of pile-v2(~1000 samples) of pile-v2 dataset, each has 1,000 random samples from the original dataset. The dataset has 255MB of text (code and english).
The dataset contains technical text on programming languages and natural language with the following subsets,
Bible
TED2020
PileOfLaw
StackExchange
GithubIssues
Opensubtitles
USPTO
S2ORC
DevDocs
CodePileReddit2022
USENET
GNOME
ASFPublicMail
PileV2Reddit2020… See the full description on the dataset page:
https://huggingface.co/datasets/CarperAI/pile-v2-small-filtered.