This dataset contains a representation of RNA sequencing data and text descriptions.
Dataset type: single (suitable for relevant contrastive-learning or inference tasks).
Cell Sentence Length: The cell sentences in this dataset have a length of $cs_length genes.
The RNA sequencing data used for training was originally gathered and annotated in the CellWhisperer project. It is derived from
CellxGene and GEO. Detailed information on the gathering and annotation of the data… See the full description on the dataset page:
https://huggingface.co/datasets/jo-mengr/human_pancreas_luecken.