KOREAN-SyntheticText is a successor of the KOREAN-WEBTEXT project in our mission to create high-quality Korean corpora. The dataset consists of 1.4B tokens generated over 600 H100 hours following the Cosmopedia project.
The dataset has been generated using a 100B + open-source LLM fine-tuned on text generation. No filtering has been done yet.