Japanese Medical VQA 12M is a large-scale Japanese medical multimodal dataset built from Open-PMC-18M and released in Parquet and Webdataset format.
This dataset contains outputs from multiple data-construction stages, including:
source captions
Japanese translations of source captions
enriched captions
Japanese translations of enriched captions
question-answering
This repository currently stores the dataset in… See the full description on the dataset page:
https://huggingface.co/datasets/MIL-UT/Japanese-Medical-VQA-12m.