A benchmark of 500 multi-choice questions for evaluating multi-hop reasoning over implicit/explicit structures extracted from scientific text. Each sample pairs a paper-grounded text passage and figure with a question, ground-truth answer, reference reasoning chain, and reference structural frame.
topic
string
Research topic… See the full description on the dataset page:
https://huggingface.co/datasets/ttwos/Anon-T2S-Bench-MR.