CCI4.0-M2 v1 is a comprehensive dataset collection consisting of two specialized subsets designed for language model training.
Notes
5.2TB Chinese webpage, 22TB English webpage, some data released in CCI4.0-M2-Extra(BAAI_datahub / modelscope / hf) due to the license concern.
430 million CoT… See the full description on the dataset page:
https://huggingface.co/datasets/BAAI/CCI4.0-M2-CoT-v1.