This package contains the tokenized Telugu pretraining dataset for the Rachana GPT project.
tokens.bin
meta.json
rachana_bpe32k.model
rachana_bpe32k.vocab
final_pretrain_corpus_v2_sangraha76.meta.json
corpus language: Telugu-first
tokenizer: SentencePiece BPE
vocab size: 32000
token count: 3,556,233,011
EOS id: 3
token dtype: uint32
This package is intended for:
causal language model… See the full description on the dataset page:
https://huggingface.co/datasets/KPrashanth/rachana-training-dataset.