This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections.
lighteval: Merge QA pairs and chunk metadata into a lighteval compatible dataset for quick model-based scoring
citation_score_filtering: Compute overlap-based citation scores and filter QA pairs accordingly
To reproduce this dataset, use YourBench v0.6.0 with the… See the full description on the dataset page:
https://huggingface.co/datasets/ChipHub/yourbench-eval-data.