Raw tokenized pretraining data for the Chess Pre-to-Post project, stored as
sharded NumPy arrays (shard_XXXX/raw.NNNN.npy).
[!IMPORTANT]
This is an earlier, smaller (20B) snapshot and is no longer maintained.
The maintained version of this dataset is
pavelslab-nyu/pretrain_v1_54B.
Please use that version for any new work — it supersedes this one.
➡️ pavelslab-nyu/pretrain_v1_54B… See the full description on the dataset page:
https://huggingface.co/datasets/chess-pre-to-post/pretrain_v1_20b.