This is a large collection of mostly synthetic instruction data in English and Finnish suitable for SFT training. For more details on our data curation and data generation pipeline, check out our Continued Pretraining Playbook.
For the Finnish portion, we translated prompts from the Tulu3 SFT Mixture into Finnish. We used Llama-3.3-70B-Instruct to generate multiple responses to the translated prompts and used the same model to select the… See the full description on the dataset page:
https://huggingface.co/datasets/LumiOpen/poro2-instruction-collection.