A refined version of StackExchange dataset in RedPajama & The Pile by Data-Juicer. Removing some "bad" samples from the original merged dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 71GB).
Number of samples: 26,309,203 (Keep ~57.89% from the… See the full description on the dataset page:
https://huggingface.co/datasets/datajuicer/redpajama-pile-stackexchange-refined-by-data-juicer.