A parallel corpus with a subset of 1200 sentence pairs of professional and laymen variants (144 019 tokens), manually simplified by linguists, as a benchmark for medical text simplification. This dataset was collected in the CLARA-MeD project, with the goal of simplifying medical texts in the Spanish language and reducing the language barrier to patient's informed decision making.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/CLARA-MeD/claramed1200.