Multiple references of raw Bengali corpus are available at this
GitHub link. Used following references from that for gathering raw bengali text for the purpose of training the tokenizer.
-
Tab-delimited Bilingual Sentence Pairs - These are selected sentence pairs from the
Tatoeba Project. This has approximately 6,500 english to bengali sentence pairs. Only Bengali sentences are extracted for training the tokenization
-
IndicParaphrase - Only the input data from validation dataset of
Bengali paraphrases are used for the tokenization. That dataset contains 10,000 Bengali sentences.