This dataset provides aligned pairs of natural language descriptions and protein
sequences, where the descriptions integrate protein family information, length
constraints, and Gene Ontology–based functional relations. The dataset is
intended for studying natural language–guided protein sequence generation and
the semantic alignment between textual protein descriptions and sequence space.
The data in this dataset are… See the full description on the dataset page: https://huggingface.co/datasets/Ethan-SHJ/NL2Protein.