Check the 🏆 leaderboard constructed with this dataset and the corresponding 👨🏻💻 evaluation code.
The SeaExam dataset aims to evaluate Large Language Models (LLMs) on a diverse set of Southeast Asian (SEA) languages including English, Chinese, Indonesian, Thai, and Vietnamese.
Our goal is to ensure a fair and consistent comparison across different LLMs on those languages while mitigating the risk of data contamination.
It consists of the following two parts:… See the full description on the dataset page:
https://huggingface.co/datasets/SeaLLMs/SeaExam.