A 50,000-row training + 1,000-row evaluation quality-filtered subset of
DKYoon/SlimPajama-6B.
Source proportions are not reweighted — the natural distribution of the
dataset is preserved. Documents are shuffled (file order + within-file) then
filtered with keep_doc() (see dataset_builders/filters.py):
Encoding
< 1 % replacement/control characters… See the full description on the dataset page:
https://huggingface.co/datasets/tutur90/slimpajama6b-50k.