Cleaned Thai text corpus assembled for pretraining a small language model
(SLM) from scratch. Built by the data_pipeline project: each upstream source
is normalized to a single text field, unicode-normalized, whitespace-cleaned,
length-filtered, and exact-deduplicated.