Dataset to accompany the paper Cost-Efficient Estimation of General Abilities Across Benchmarks.
WILD-raw contains the full evaluation responses for 65 language models across 27 benchmarks (109,566 unique items), including conversations, model answers, targets, and scorer output.
For a lightweight version with just scores and token usage, see WILD.
7,237,945 total (model, item) observations
65 models… See the full description on the dataset page:
https://huggingface.co/datasets/kensho/WILD-raw.