Human-validation subset of the SegVQA Kvasir-SEG benchmark
(segvqa_benchmark_kvasir.json), 100 verified items sampled with seed 42,
stratified by category (multi-hop/comparison slightly oversampled).
Q — is the question sensible and answerable… See the full description on the dataset page:
https://huggingface.co/datasets/AiventraLab/segvqa-review.