The same 200,913 rows as
ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80, rewritten in one
random order so that a prefix is a sample.
The source repo is sorted by protein accession. That makes the first N rows a
block of related proteins rather than a sample of the pool, so any subset that
is actually representative requires reading all 154.01 GB and
selecting rows afterwards. Here the rows are stored shuffled, so… See the full description on the dataset page:
https://huggingface.co/datasets/ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80-shuffled.