This dataset consists of 120 million texts (approximately 89.3B tokens) filtered from the 376 million Japanese texts in FineWeb2 that were deemed educational. The following subsets are also provided:
default: Approximately 120M texts (120 million texts) totaling around 89.3B tokens
sample_10BT: A random sample of about 10B tokens from the default dataset
small_tokens: Data composed solely of texts with 512 tokens… See the full description on the dataset page:
https://huggingface.co/datasets/hotchpotch/fineweb-2-edu-japanese.