A multi-task evaluation benchmark for large language models on the Mongolian language (Cyrillic script). Six task configurations cover open-ended QA, multiple-choice, code generation, instruction following, math, and culturally grounded knowledge.
math
150
Numeric / short answer
prompt, answer, accepted_formats, rationale… See the full description on the dataset page:
https://huggingface.co/datasets/Bokhbat/LLM_Benchmark.