A federated learning-ready version of the PathVQA dataset, partitioned by image using perceptual hashing (pHash). Each client's data contains all question-answer pairs linked to unique images, mimicking a real-world distributed medical setting (e.g., hospitals with private slide collections).