This dataset is part of the research paper "Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai" submitted to ACL SRW 2024. It represents the best-performing synthetic dataset (F+C+D+) generated using our novel seed-free framework for low-resource languages, specifically Thai.
Size: 5,000 instructions
Language: Thai
Task: Instruction-tuning for Large Language Models… See the full description on the dataset page:
https://huggingface.co/datasets/parinzee/seed-free-synthetic-instruct-thai-v1.