This repository contains pre-tokenized data used in the paper Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws.
The dataset consists of pre-tokenized .pt files containing packed GPT-2-tokenized sequences. These snapshots were used to study the data-constrained, compute-rich regime of pretraining where models are trained for multiple epochs over a finite dataset.