ROOTS Subset: roots_id_indo4b_jw300
Dataset uid: indo4b_jw300
Indo4B consists of around 4B words, with around 250M sentences. The dataset covers both formal and colloquial Indonesian sentences compiled from 12 corpus, of which two corpus cover Indonesian colloquial language, eight corpus cover formal Indonesian language, and the rest have a mixed style, both colloquial and formal.\n