An open-source pretraining dataset containing 4690 billion tokens, this bilingual dataset with both English and Chinese texts is used for training neo models.
The dataset consists of several components, each originating from different sources and serving various purposes in language modeling and processing. Below is a brief overview of each component:
Common Crawl
Extracts from the Common Crawl project, featuring a rich diversity of… See the full description on the dataset page:
https://huggingface.co/datasets/m-a-p/Matrix.