A large-scale corpus of Standard Classical Tibetan text prepared specifically for training character-level KenLM n-gram language models for use in the PaganTibet normalisation pipeline. The dataset contains approximately 17.7 million lines of cleaned, line-split ACTib text in two versions: non-tokenised and tokenised (via a customised version of the Botok Tibetan tokeniser). These two files are the direct training corpora for… See the full description on the dataset page:
https://huggingface.co/datasets/pagantibet/Tibetan-KenLM-ACTibtrainingdata.