This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code:
https://github.com/opendatalab/Meta-rater
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page:
https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.