Tajik Linguistic Benchmark (TajLiB)
Dataset Description
TajLiB (Tajik Linguistic Benchmark) is a comprehensive multiple-choice question benchmark designed to evaluate the linguistic understanding of Large Language Models (LLMs) in the Tajik language.
It assesses the model's knowledge of Tajik phonetics, grammar, lexicology, and orthography, as well as its ability to understand and reason about domains related to Tajik language usage, including legal texts and Tajik… See the full description on the dataset page: https://huggingface.co/datasets/f1rdavs/TajLiB.