A 278k sample derivation of the first 3M samples from the C4 dataset for a cheap and short continued pretraining for language models to optimize for benchmark scores without sacrificing generalization and generative modelling unrelated to chat or 'instruct' data.
The estimated top 10% of highest estimated length normalized ngram (mean of tri, quad, and penta-gram) overlaps for each of the
selected benchmark datasets (arc, truthful_qa, hellaswag, mmlu… See the full description on the dataset page:
https://huggingface.co/datasets/crumb/c4-benchfilter-nano.