The dataset is sourced from van der Flier et al. (2024), available at
https://www.sciencedirect.com/science/article/pii/S2001037024002940. It contains a total of 3706 rows, with each sample carrying 1 to 8 mutation sites. The full dataset is randomly split into training, validation, and test sets following an 8:1:1 ratio. The target label is absorbance, ranging from -0.001 to 0.211. A higher absorbance value indicates greater starch degradation, corresponding to stronger detergent activity of the amylase enzyme.
Performance (on test set)