2023.12.20更新:增加来自skypile数据集的长数据
Long text dataset for pretraining
This dataset contains samples with the length greater than 16k, which can be used for pretraining models with extremely long context lengths.
The dataset is continuously updating.
此数据集包含长度大于16k的预训练数据,可用于对极长上下文长度的模型进行预训练。
数据正在持续增加中
筛选自 悟道200G开源数据、书生万卷数据集、
CCI中文互联网语料库
、中文维基百科等,
每条数据长度在16000字以上
筛选自 [SlimPajama-dc]… See the full description on the dataset page:
https://huggingface.co/datasets/yuyijiong/LongData-Corpus.