100,000 documents from
mlfoundations/dclm-baseline-1.0
(CC-BY-4.0), each labeled with its training influence toward 17
held-out text categories: the cosine between the document's training
gradient and each category's mean held-out gradient, measured at the
end-of-stable-LR checkpoint of a 56M reference model (PleIAs/Monad 8k
tokenizer, 512 hidden × 12 layers, ~0.9B dclm tokens).
These… See the full description on the dataset page:
https://huggingface.co/datasets/Lambent/influence-taste-labels.