nanochat の日本語フォーク nanochat-jp で使用する 事前学習用日本語コーパス です.
LLM によるクリーニングを施した日本語ウェブテキストと,llm-jp の公開コーパスを混合したものを,nanochat のデータローダがそのまま読める parquet 形式で配布しています.
llm-jp/scaling-data-constrained-llms… See the full description on the dataset page:
https://huggingface.co/datasets/tohoku-nlp/nanochat-jp-pretrain.