Repository: skyfalconai/MultiLingual-Global-CorpusSize: ~6GB of plain textLanguages: 40Use Case: Training multilingual or Indic language tokenizers and language models
MultiLingual-Global-Corpus is a diverse and curated collection of multilingual and Indic language text data, sourced from a blend of open datasets including:
AI4Bharat Shrutilipi Voice corpora
OSCAR multilingual datasets
User-contributed and scraped text corpora… See the full description on the dataset page:
https://huggingface.co/datasets/skyfalconai/MultiLingual-Global-Corpus.