OpenSakura-DS-260220-LN-ja-zh-PT-Adam is a large-scale pretraining corpus built from light-novel source shards and filtered to Japanese/Chinese scripts.
This export is intended for PT/CPT-style language modeling.
Dataset Summary
Metric
Value
Dataset ID
OpenSakura/OpenSakura-DS-260220-LN-ja-zh-PT-Adam
Total rows
9,515,512
Total parquet files
480
Total size
63,621,025,693 bytes (~63.62 GB, ~59.25 GiB)