See our preprint for details and our GitHub repo for pretraining, finetuning and benchmarking scripts.
Join our slack community for support and discussion about microbiome foundation models.
Each row is one sample. Taxa and Relative Abundances are aligned lists (same length per row): semicolon-separated rank strings for taxa, and matching relative abundances. Additional columns record data type, sequencing method, pipeline version, and study… See the full description on the dataset page:
https://huggingface.co/datasets/outpost-bio/Atlas.