This dataset contains 16,384 base pair segments from the human reference genome (GRCh38) prepared for Sparse Autoencoder (SAE) training with the Evo2 model. The segments are extracted using a sliding window approach with 75% overlap.
Dataset Details
Total segments: 718,648
Segment size: 16,384 base pairs
Stride: 4,096 base pairs (75% overlap)
Source genome: GRCh38.primary_assembly (GENCODE Release 41)… See the full description on the dataset page: https://huggingface.co/datasets/harari/human_grch38_segment_sample.