This is a 100.0 Million token subset of krisbailey/RedPajama-Data-V2-1B, which is a subset of togethercomputer/RedPajama-Data-V2.
Motivation
100M tokens is a standard size for:
CI/CD Pipelines: Fast enough to download and train for unit tests.
Debugging: Verifying training loops without waiting for hours.
Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B).
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-100M.