A multi-task NLP dataset in the Kazakh language, covering seven distinct language tasks - from instruction-following and question answering to translation, sentiment analysis, and grammar exercises. Designed to support the development of Kazakh-language models, benchmarks, and linguistic research.
Kazakh is a Turkic language spoken by over 13 million people, yet it remains significantly underrepresented in NLP research and… See the full description on the dataset page:
https://huggingface.co/datasets/mangi-llm/Kazakh_Multi-Task_corpus.