To train Med-V1, we construct MedFact-Synth, a large-scale synthetic training set including 1.5 million instances.
Each instance contains: a synthetic claim to be verified, a source article serving as evidence, a rationale explaining the verification, and a 5-point Likert-scale verdict, ranging from strong contradiction (-2) and partial contradiction (-1) to neutral (0), partial agreement (+1), and strong agreement (+2).
To build this dataset, we begin by sampling one million articles from… See the full description on the dataset page:
https://huggingface.co/datasets/ncbi/MedFact-Synth.