A canonical 10 Billion token weighted subset of the RedPajama-Data-1T dataset.
This dataset is a faithful reproduction of the original RedPajama-Data-1T distribution, scaled down to exactly 10 Billion tokens. It is designed to preserve the exact domain ratios of the original dataset (excluding the defunct 'Books' subset). This allows researchers and developers to prototype, debug, and test on a representative slice of the data… See the full description on the dataset page:
https://huggingface.co/datasets/krisbailey/RedPajama-10B-Weighted.