SEA-LION-Pile is the pretraining data set for SEA-LION, a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region.
This repository contains the cleaned mC4 portion of the SEA-LION-Pile.
For the remainder of the SEA-LION-Pile dataset, they may be downloaded from the links provided below.