Can a protein language model recognise that two proteins share a fold when their
sequences look unrelated? That is the question this dataset asks.
You are given a large lookup set of protein domains with known CATH
structural classifications, and a small set of query domains that have been
filtered so no query has a detectable sequence-alignment relative in the lookup
set. Predict each query's CATH class by finding its… See the full description on the dataset page:
https://huggingface.co/datasets/GrimSqueaker/cath43-eat.