Language: Maithili (मैथिली, ISO 639-3: mai)
License: CC-BY-4.
Size: 28,622 documents | 13.7M words | ~45.8M tokens
Format: JSONL (one paragraph per row)
Tags: unlabelled, low-resource, indic-nlp, monolingual, pretraining
A large, unlabelled corpus of written Maithili text for language model pretraining and unsupervised NLP research.
Property
Value
Documents
28,622
Total Words
13,664,375
Total Subword Tokens… See the full description on the dataset page:
https://huggingface.co/datasets/kamal-018/Maithili-Corpus.