B-CORE (Bengali Context-aware Optimized and Refined Entities) is a large-scale, rigorously curated Bangla monolingual corpus for language model pretraining, comprising 16.5 million documents (4.32 billion tokens, 52GB (20.8 GB Compressed)). It is among the largest and most carefully curated Bangla pretraining corpora available, constructed through a reproducible multi-stage pipeline.
B-CORE was used to pretrain the BnLM-F and BnLM-C Bengali… See the full description on the dataset page:
https://huggingface.co/datasets/nahid-hub/B-CORE-bengali-corpus.