The dataset SToCorpus-88M is the pre-training dataset used by the SToFM model.
Paper:
https://arxiv.org/abs/2507.11588
SToFM Model Github:
https://github.com/PharMolix/SToFM
If you find SToCorpus-88M helpful to your research, please consider giving our Github repository a 🌟star and 📎citing the following article. Thank you for your support!
@article{zhao2025stofm,
title={SToFM: a Multi-scale Foundation Model for Spatial Transcriptomics},
author={Zhao, Suyuan and Luo… See the full description on the dataset page:
https://huggingface.co/datasets/Toycat/SToCorpus-88M.