Occurrences of every probed entity name in the OLMo-2 PRETRAINING corpus (olmo-mix-1124, 1,117 token files, 14.10 TiB), counted as OLMo token sequences with a word boundary required at each end, both written forms kept separate. SUPERSEDES factprobe-replication-stage1-counts-olmotok-v1, which was missing each entity's canonical name: the released triples name entities by their Wikidata aliases, a field that by construction… See the full description on the dataset page:
https://huggingface.co/datasets/latkes/factprobe-replication-stage1-counts-canonical-v1.