A vision–language dataset of peroxisome-biology figures paired with their published captions — for training and evaluating multimodal models on PEX10 and peroxisome-biogenesis-disorder literature.
Figure ↔ caption pairs in LLaVA format, extracted from PEX10-related PubMed Central Open Access (CC-BY) articles. 2,985 figure images become ~9,700 multimodal training records across four complementary task formats — plain captioning… See the full description on the dataset page:
https://huggingface.co/datasets/SkyWhal3/PEX10_PubMed_Central_MultiModal.