Pre-tokenized corpora for midtraining intro-timing experiments: studying how the point
in a pretraining schedule at which you introduce a new data mixture (code, math, chemistry)
changes the final model.
These are not new corpora. Every file is a token-ID encoding of an existing public
dataset, published so the experiments are reproducible without re-running tokenization
(which is the expensive part). All sources are credited… See the full description on the dataset page:
https://huggingface.co/datasets/Impliedhomeland/midtrain-bridge-data.