Raw stage-1 pre-pretraining corpus from the ppt
research framework, exported for information-theoretic analysis (e.g.
m-local entropy) independent of this repo's training pipeline.
Uniform random digits in [0, 10), sampled independently per position — a structureless baseline.
These input_ids are not decodable with the Pythia (or any other)
tokenizer. They are integers in [0, 10) with… See the full description on the dataset page:
https://huggingface.co/datasets/sashaboguraev/ppt-random_numbers-corpus.