A benchmark Language Modeling dataset for Tagalog. The dataset construction was done similar to that of the WikiText
Long Term Dependency Language Modeling Dataset, with a some differences, such as in how Wikipedia was scraped and how the vocabulary was
created. The dataset contains 39 Million tokens in the training set.