We assembled a large-scale pretraining corpus, Genecorpus-104M, comprised of ~104 million human single cell transcriptomes from a broad range of tissues from publicly available data. This corpus was used for pretraining Geneformer-V2, a pretrained transformer model that enables context-aware predictions in settings with limited data… See the full description on the dataset page: https://huggingface.co/datasets/theodoris-lab/Genecorpus-104M.