Per-generation records from seven language models, capturing the token IDs each
model actually emitted alongside the canonical re-encoding of its own output —
plus the trained toy-model checkpoints from the accompanying controlled
experiment.
Code, writeup and full run log:
https://github.com/brendanlong/tokenization-hidden-computation-experiment
Tokenization is many-to-one: many token sequences decode to the same string, but… See the full description on the dataset page:
https://huggingface.co/datasets/brendanlong/retok-noncanonical-tokenization.