The Datasets contains the jargon terms, lay definitions, general definitions for different stages in our REAME pipeline. To comply with fair use of law~\footnote{\url{
https://www.copyright.gov/fair-use/}}, We used GPT-3.5 to paraphrase the lay definitions as shown in synthetic_data_creation.ipynb. We used GPT-4o-mini to paraphrase the EHRs as shown in synthetic_EHR_creation.ipynb. We asked the LLM(gpt-4o-mimi) to edit the original sentence but make sure to keep the main terms unchanged. Check… See the full description on the dataset page:
https://huggingface.co/datasets/bio-nlp-umass/NoteAid-README.