Collective-Corpus is a massive-scale, multi-domain dataset designed to train Transformer-based language models from scratch and finetune them across a wide variety of domains — all in one place.
This dataset aims to cover the full LLM lifecycle, from raw pretraining to domain-specialized finetuning.
Large-scale, diverse multilingual text… See the full description on the dataset page:
https://huggingface.co/datasets/dignity045/Collective-Corpus.