Humans conserved cCRE enhancers (v18) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
752,548 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page:
https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-humans-intervals-v18_255_128.