500 student–tutor dialogues selected and normalized from a larger raw pool of
~1,900 synthetically generated dialogues (source files: exams,
maths_and_informatics, mixed_themes, physics_and_informatics), each of
which originally used a different JSON schema. This file merges them all
into one consistent schema, removes duplicates and broken records, and
selects a maximally diverse subset for LoRA/SFT fine-tuning of a small
(1.5B)… See the full description on the dataset page:
https://huggingface.co/datasets/ptvnck/TutoringDialogs.