This dataset provides a benchmark for evaluating the scalability of genomic models to even-longer DNA inputs through a mutation hotspot classification task. Using whole-genome variant data from the Chinese Pangenome Consortium (CPC) (Gao et al., 2023), we identify genomic regions (hotspots) exhibiting significantly higher mutation densities compared to local chromosomal backgrounds. Sequences of 8 Kbp, 32 Kbp, and 128 Kbp are extracted to create three parallel tasks… See the full description on the dataset page:
https://huggingface.co/datasets/BGI-HangzhouAI/Benchmark_Dataset-variant_hotspot.