FreshQA Kazakh is a machine-translated version of the original FreshQA hallucination benchmark dataset. It is designed to test the factuality and hallucination rates of large language models, with a specific focus on questions whose answers may change over time or involve false premises.
This dataset contains exactly 600 samples categorized into three distinct fact types (fact_type). To account for various ways a model might correctly… See the full description on the dataset page:
https://huggingface.co/datasets/issai/freshqa_kazakh.