FactRBench is a benchmark designed to evaluate the factuality of long-form responses generated by large language models (LLMs), focusing on both precision and recall. It is released alongside the paper [VERIFACT: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Facts].
Current factuality evaluation methods emphasize precision—ensuring statements are… See the full description on the dataset page:
https://huggingface.co/datasets/launch/FactRBench.