CS8 rand dataset — random inner alias assignment, negative control for mechanistic interpretability study
train.jsonl — 12,000 training examples (20 aliases × 600)
val.jsonl — 12,000 validation examples (20 aliases × 600)
alias_vocab.json — 20 aliases with T1/T2 tokenization group labels
T1 (single-token): emp, inv, txn, mgr, ord, prod, cust, dept, acct, sale
T2 (two-token, generic first subtoken): shp, whs, rgn… See the full description on the dataset page:
https://huggingface.co/datasets/Likithp/cs8-rand.