A cleaned, split, and leakage-controlled build of PubMedQA
(Jin et al., 2019) for fine-tuning a generative LLM to read a PubMed abstract plus a research question and
answer yes / no / maybe, with a deliberate focus on calibration and the minority maybe class.
The data is stored as raw text (question, abstract, label) — no tokenization, no rendered chat
prompt — so it is model-agnostic. Different base models (e.g. Llama-3.1-8B… See the full description on the dataset page:
https://huggingface.co/datasets/Legeng/pubmedqa-prepared.