Here we provide pruned TCGA transcriptomics data from manuscript "MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention". Code is available at GitHub.
The TCGA [1] transcriptomics data were collected from Xena [2] and preprocessed using the proposed novel pipeline in MIRROR [3].
For raw transcriptomics data, we first apply RFE [4] with 5-fold cross-validation for each cohort to identify the most performant support set for the subtyping… See the full description on the dataset page:
https://huggingface.co/datasets/Franklin2001/MIRROR_Pruned_TCGA_RNASeq_Data.