This dataset is part of PepBenchmark, a standardized benchmark for peptide machine learning introduced in the paper PepBenchmark: A Standardized Benchmark for Peptide Machine Learning.
PepBenchmark unifies datasets, preprocessing, and evaluation protocols for peptide drug discovery. It comprises three components:
PepBenchData: A well-curated collection of 29 canonical-peptide and 6 non-canonical-peptide datasets across 7 groups.
PepBenchPipeline: A standardized preprocessing pipeline… See the full description on the dataset page:
https://huggingface.co/datasets/yisen888/PepBenchData.