GigaKriya: A Large Bengali Text Corpus with Educational and Toxicity Annotations
Dataset Summary
This repository contains a large corpus of Bengali text (~25 billion tokens), which has been filtered and annotated using classifiers for educational content and toxicity. The dataset is intended for training language models and other NLP applications in Bengali. GigaKriya is part of the Polyglot project, which aims to develop multilingual resources and models for… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/gigakriya-v1.