Benchmark for protein sequence alignment under remote homology.
Subsets
sup: low-to-intermediate identity
twi: very low identity (<25%)
sup_fp: superfamilies with added false positives
twi_fp: twilight with added false positives
all: union of the above
Features
pair_id, group_id, set_name
seq1_id, seq2_id, seq1, seq2
ref_alignment: list of index pairs [[i, j], ...] (0-based)
percent_identity, scop_labels, meta
Usage… See the full description on the dataset page: https://huggingface.co/datasets/DeepFoldProtein/SABmark-dataset.