Serendip-SFT-Sinhala is a large-scale Sinhala instruction-tuning dataset with 293,613 high-quality examples for supervised fine-tuning (SFT) of large language models.
Created to train SerendipLLM, a Sinhala language model designed to excel at instruction-following, question-answering, summarization, and text classification.