Leakage-free protein function prediction benchmarks for multi-task
protein language model (pLM) fine-tuning.
This dataset contains 16 prediction tasks drawn from
three independent data families, resplit to eliminate cross-dataset
sequence similarity-based data leakage between training, validation,
and test sets.
secstr
per-residue classification
3
accuracy
DSSP assignments from… See the full description on the dataset page:
https://huggingface.co/datasets/Moomboh/mutafitup-datasets.