This repository provides the dataset resources used for training and evaluating SlideChat, a multimodal large language model for whole-slide pathology image understanding.
The dataset includes both instruction-following training data and VQA/Caption evaluation benchmarks across multiple pathology cohorts and tasks.
SlideInstruct_train_stage1_caption.json: Slide-level caption instruction data used for Stage-1 training.… See the full description on the dataset page:
https://huggingface.co/datasets/General-Medical-AI/SlideChat.