Villanova-SFT-2603 is a large-scale, multilingual supervised fine-tuning (SFT) collection of datasets. It contains 1,711,114 instruction-response conversations spanning five European languages, covering chat, instruction following, reasoning, code, knowledge, and safety tasks. This dataset was used to train the Villanova-2B-2603 model family.
All data has been processed through a rigorous curation pipeline that enforces schema normalization, hash-based… See the full description on the dataset page:
https://huggingface.co/datasets/VillanovaAI/villanova-sft-2603.