DMS Benchmark includes three types of mutations: indels, single substitution and multiple substitutions.
Each tsv file contains all data for a single task, representing all possible mutations on a specific protein sequence. The filename corresponds to the task name. Here we use randomly cross-validation scheme, where the data for each task is randomly divided into five folds and the fold_id column indicates the fold assignment. The labels are continuous… See the full description on the dataset page:
https://huggingface.co/datasets/genbio-ai/ProteinGYM-DMS.