A small, high-quality supervised fine-tuning (SFT) dataset of 559 Tamil-language
medical instruction–reasoning–answer triples. Every item is written in natural Tamil
script, shows explicit step-by-step clinical reasoning before the answer, and was
independently audited by an LLM judge and decontaminated against public Tamil
medical benchmarks.
The goal is to teach a model how a clinician reasons in Tamil — across frontline care,
basic-science mechanisms, and… See the full description on the dataset page:
https://huggingface.co/datasets/Sidharth1743/tamil-med-sft-cot-600.