This dataset contains data for individual evaluation instances
from the DataDecide project (publication forthcoming). It shows how
standard evaluation benchmarks can vary across many dimensions of
model design.
The dataset contains evaluations for a range of OLMo-style models
trained with:
25 different training data configurations
9 different sizes with parameter counts 4M, 20M, 60M, 90M, 150M, 300M, 750M, and 1B
3 initial random seeds
Multiple… See the full description on the dataset page:
https://huggingface.co/datasets/allenai/DataDecide-eval-instances.