A curated Slovenian medical instruction corpus for supervised fine-tuning and research
in Slovenian biomedical NLP. Machine-translated from open English medical datasets into
formal Slovenian medical register, with a strict term-preservation gate (numbers, units,
drug names, Latin terms and glossary terms must survive translation, else the row is dropped).
Research / educational use only. Not for clinical use. These are… See the full description on the dataset page:
https://huggingface.co/datasets/texdata/med-slo-sft.