This dataset is the reupload of ai4bharat/sangraha dataset. Specifically, 1.9 Million rows of Hindi Verified Data. This is tokenized with Hindi Tokenizer: atharvanighot/hindi-tokenizer such that it can be used to train directly as it is pretokenized dataset.