We have released a portion of the sampled 100B tokens data from the CCI4.0-M2 v1, including Chinese and English Web datasets, domain-specific datasets and Chain-of-Thought reasoning datasets.
The dataset directory is consistent with the overall dataset. For a detailed description of the dataset, please refer to CCI4.0-M2 v1 README.
Name
Tokens
Tokens(B)… See the full description on the dataset page:
https://huggingface.co/datasets/BAAI/OpenSeek-Pretrain-100B.