Raw stage-1 pre-pretraining corpus from the ppt
research framework, exported for information-theoretic analysis (e.g.
m-local entropy) independent of this repo's training pipeline.
Neural Cellular Automata grid-rollout patch IDs (positional-numeral encoding of 2x2 color patches, plus start/end tokens), generated by submodules/nca-pre-pretraining.
These input_ids are not decodable with the Pythia (or… See the full description on the dataset page:
https://huggingface.co/datasets/sashaboguraev/ppt-nca-corpus.