The Manipuri Wikipedia Corpus is a pure Manipuri text dataset derived from the Manipuri-language Wikipedia (as.wikipedia.org).
It contains cleaned plain text extracted from Wikipedia articles, stripped of all formatting, with non-Manipuri characters completely removed.
This dataset is designed for language modeling, NLP research, creating Manipuri specific tokenizers, and other Manipuri-language processing tasks.… See the full description on the dataset page: https://huggingface.co/datasets/marsh-mellow/manipuri_wikipedia.