Full evaluation traces across Corral environments, models, agents, and task granularities
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the full evaluation traces collected across all 8 Corral environments.
Each configuration (config) corresponds to a unique combination of model, environment, scope (difficulty… See the full description on the dataset page:
https://huggingface.co/datasets/jablonkagroup/corral-traces.