The PileV2 is a larger and more diverse collection of text data mostly focused on English text. Specifically, it is a collection of roughly 40 different data subsets. This includes the original 22 subsets from the original pile plus a heavy focus on additional software engineering specific data subsets including the newly released "The Stack" from bigcode, various programming competition sources, and a number of programmer… See the full description on the dataset page: https://huggingface.co/datasets/CarperAI/pilev2-dev.