Paper: QuRating: Selecting High-Quality Data for Training Language Models
A 260B token subset of cerebras/SlimPajama-627B, annotated by princeton-nlp/QuRater-1.3B with sequence-level quality ratings across 4 criteria:
Educational Value - e.g. the text includes clear explanations, step-by-step reasoning, or questions and answers
Facts & Trivia - how much factual and trivia knowledge the text contains, where specific facts and obscure trivia are preferred over more… See the full description on the dataset page:
https://huggingface.co/datasets/princeton-nlp/QuRatedPajama-260B.