The data in this dataset has the promoter sequences and the corresponding gene expression data as TPM values for 26 Maize NAM lines and has been used for the finetuning step (for the downstream task of gene expression prediction) of Florabert.
The data has been split into train, test and eval data (70-20-10 split). In all, there are ~ 7,00,000 entries across the files. The steps followed to obtain this… See the full description on the dataset page:
https://huggingface.co/datasets/Gurveer05/maize-nam-gene-expression-data.