This is our SEA translationese vs. natural classification dataset for the "SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages" paper.
SEACrowd is a collaborative initiative that consolidates a comprehensive resource hub that fills the resource gap by providing standardized corpora in nearly 1,000 Southeast Asian (SEA) languages across three modalities.
To analyze the generation quality of LLMs in SEA languages, we build a text… See the full description on the dataset page:
https://huggingface.co/datasets/SEACrowd/sea_translationese_resampled.