Averaged QA results for factual-knowledge and reasoning evaluations across all 8 Corral environments
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the averaged results of the question-answer evaluations used to test the factual knowledge and reasoning ability of models across all 8 Corral environments.
The… See the full description on the dataset page:
https://huggingface.co/datasets/jablonkagroup/corral-QAs-topic_reports.