SEA-Instruct-2602 is a preliminary release of instruction-tuning data focused on Southeast Asian languages and contexts. The dataset combines prompts filtered from open-source data with our own synthetic prompts, paired with synthetic responses, for language model training on SEA-specific tasks and languages.
This dataset contains only the filtered subset of data with prompt_input_quality at Excellent, prompt_is_coherent at True and… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/SEA-Instruct-2602.