This dataset is the Kyrgyz-translated version of the GSM8K benchmark, which evaluates a model's ability to solve grade-school math word problems.
This dataset is a component of the KyrgyzLLM-Bench, a comprehensive suite for evaluating LLMs in Kyrgyz.
Main Paper: Bridging the Gap in Less-Resourced Languages: Building a Benchmark for Kyrgyz Language Models
Hugging Face Hub:
https://huggingface.co/TTimur
GitHub Project:… See the full description on the dataset page:
https://huggingface.co/datasets/TTimur/gsm8k_kg.