A dataset of over 1.3M bacterial genome embeddings together with associated metadata.
The dataset was collated by combining publicly available bacterial genomes from MGnify, SPIRE and
NCBI. The extracted genomes were then embedded with a pretrained Bacformer model (masked objective).
The genome embeddings were calculated by taking the average of contextual protein embeddings from the last layer of the Bacformer encoder.… See the full description on the dataset page:
https://huggingface.co/datasets/macwiatrak/bacformer-genome-embeddings-corpus.