The Bambara-Texts dataset is a collection of monolingual Bambara text designed for pretraining language models. It provides a diverse set of textual data to improve natural language processing (NLP) applications for the Bambara language.
This dataset can be used for:
Pretraining large language models (LLMs)
Building word embeddings for Bambara
Language modeling tasks such as masked language modeling (MLM) and autoregressive modeling
Corpus-based linguistic research… See the full description on the dataset page:
https://huggingface.co/datasets/djelia/bambara-texts.