Dataset Card for BSC Multilingual Synthetic SFT Instructions
Dataset Summary
This dataset consists of 714k conversations mixing human and synthetic instructions generated to post-train language models across five languages: Catalan, Spanish, English, Basque, and Galician.
This dataset has been used in the supervised fine-tuning stage of ALIA-40b-instruct-2606. The training mixture is obtained by combining a selection of (human and synthetic) permissively licensed… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-SFT.