Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments.
The dataset is organized into 24 configurations… See the full description on the dataset page:
https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.