A community database of AI evaluation results, all in one schema. Scores scraped from
leaderboards, pulled out of papers, and produced by local evaluation runs are stored in a
single record format, so results from different sources can be compared, joined, and reused
instead of re-scraped. This dataset is the data itself: one JSON record per
model per evaluation run — which may carry several scored results — with optional
per-sample companion files.… See the full description on the dataset page:
https://huggingface.co/datasets/evaleval/EEE_datastore.