This dataset provides a benchmark for evaluating the model's ability to leverage richer genetic information from longer sequences to achieve more accurate inference.
Using data from the Human Pangenome Reference Consortium (BioProject ID: PRJNA730823), we designed a population classification task focusing on African, East Asian, and European population groups.
From samples' VCF file and the reference genome sequence, we generated sample pseudo-sequences.
Based on variant… See the full description on the dataset page:
https://huggingface.co/datasets/BGI-HangzhouAI/Benchmark_Dataset-Human_population_classification.