Precomputed GSVA scores for bulk RNA-seq and pseudobulk AnnData files. Each
top-level directory is named after its original .h5ad file (without the
extension).
All samples are scored against the same frozen empirical reference, built from
714,800 non-benchmark ARCHS4 and pseudobulk samples across 19,260 genes.
Expression is transformed with CP10k + log1p. For every target sample, each
gene is mapped to its percentile in the frozen per-gene reference… See the full description on the dataset page:
https://huggingface.co/datasets/rpowalski/gsva_global.