OpenGenome2 is a database of nearly 9 trillion base pairs of curated DNA from across all domains of life. Collected from diverse species and public data sources, OpenGenome2 was used to train Evo 2 models. Please refer to the Evo 2 preprint or github repository for further details and usage examples.
We provide OpenGenome2 in two formats, the dataset is organized into two main directories to reflect this:
fasta which contain the DNA sequences
jsonl which include… See the full description on the dataset page:
https://huggingface.co/datasets/satputekuldip/opengenome2.