The Protein2Text-QA dataset is designed to generate human-readable explanations for protein functions based on protein sequences. It consists of question-answer (QA) pairs generated from PubMed Central (PMC) articles using LLaMA3.1-8B-Instruct. The dataset is structured into different subsets tailored for pretraining, fine-tuning, and evaluation.
Size: ~210,000 QA pairs
Source: UniProt (pretraining), PubMed Central (PMC) (QA… See the full description on the dataset page:
https://huggingface.co/datasets/tumorailab/Protein2Text-QA.