Sea-bench is a multilingual benchmark for assistant-style models annotated by native linguists
covering 8 Southeast Asian languages. The linguists sourced such data by manually translating
open-source English test sets, collecting real user questions from local forums and websites,
collecting real math and reasoning questions from reputable sources, as well as writing test
instructions and questions themselves. The Sea-bench test set contains 20 questions per task
(5 tasks for 3 languages, 4 tasks for other 5 languages).