Selected protein similarities within training, test, and validation sets.
Each protein gets two similarities selected at random and (usually) proteins within the top and bottom quintiles for similarity.
The protein is represented by its UniProt ID and its amino acid sequence (using IUPAC-IUB codes where each amino acid maps to a letter of the alphabet, see:
https://en.wikipedia.org/wiki/FASTA_format ).
The distance column is cosine distance (identical =… See the full description on the dataset page:
https://huggingface.co/datasets/monsoon-nlp/protein-pairs-uniprot-swissprot.