This is the dataset used for pre-trained plant foundation DNA large language models.
The dataset contains a collection of 22 plant reference genomes. Genome sequences are processed to fit lengths range from 1 bp to 2000 bp.Both hardmasked genomes and unmasked genomes are used to generate the pre-train dataset.