The OMG is a 3.1T base pair metagenomic pretraining dataset, combining EMBL's MGnify and JGI's IMG databases. The combined data is pre-processed into a mixed-modality dataset, with translated amino acids for protein coding sequences, and nucleic acids for intergenic sequences.
We make two additional datasets available on the HuggingFace Hub:
OG: A subset of OMG consisting of high quality genomes with taxonomic information.
OMG_prot50: A protein-only… See the full description on the dataset page:
https://huggingface.co/datasets/tattabio/OMG.