This dataset is a filtered version of the large-scale Japanese pre-training corpus, abeja-cc-ja. It was created by duplicating the original dataset while excluding texts with a KenLM perplexity score. The goal is to extract high-quality Japanese text data to facilitate the training of language models.
Data Structure
Features
Each data entry contains the following attributes: