Long-form, item-level response data harvested from 200+ AI evaluation benchmarks, standardized into a single registry-backed schema for Item Response Theory (IRT) / psychometric analysis of AI systems. This is a project of the AIMS Foundation and feeds the torch-measure toolkit.
Unlike a leaderboard (one aggregate score per model per benchmark), our data bank keeps every individual (subject, item) observation, including the model that answered, the item it… See the full description on the dataset page:
https://huggingface.co/datasets/aims-foundations/measurement-db-archived.