Thai-English Corpus is a large-scale bilingual corpus containing Thai and English text collected from publicly available datasets on Hugging Face.
The corpus combines educational content, web documents, Wikipedia articles, legal texts, medical articles, financial documents, software-related text, and other general-domain content into a unified format suitable for language model pretraining and NLP research.
Each document is stored with its original source… See the full description on the dataset page:
https://huggingface.co/datasets/Naphon/Thai-English-Corpus.