IndicCorp v2 Dataset
Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages
This repository contains the pretraining data for the paper published at ACL 2023.
dataset = load_dataset("ai4bharat/IndicCorpV2", "indiccorp_v2", data_dir="data/tel_Telu")
All the datasets created as part of this work… See the full description on the dataset page:
https://huggingface.co/datasets/bindaas200500/IndicCorpV2.