Pretraining corpus for the LoRACLE — a weight-reading interpretability model
that describes what a LoRA adapter was trained on by reading its direction
tokens. Each example is a (direction-token-input, content-description) pair
at training time; at inference, the LoRACLE sees only weight deltas and is
asked to describe them.
Split
Rows
Organisms
Toxic rows
train
50,000
25,000
2482 (5.0%)
dpo_heldout
500
250
32
val
100
50
6… See the full description on the dataset page:
https://huggingface.co/datasets/ceselder/loracle-pretrain-mix.