The bible-en-th dataset is a bilingual corpus containing English and Thai translations of the Bible, specifically the King James Version (KJV) translated into Thai. This dataset is designed for various natural language processing tasks, including translation, language modeling, and text analysis.
Languages: English (en) and Thai (th)
Total Rows: 31,102
The dataset consists of two main features:
en: English text from… See the full description on the dataset page:
https://huggingface.co/datasets/Tsunnami/bible-en-th.