Hugging Face dataset repository: evaluation-ecosystem/evaluation-ecosystem-data.
Simulation outputs supporting the AI Evaluation Ecosystem paper. Each run is a stochastic
simulation of an AI evaluation ecosystem (providers, evaluators, consumers, regulators,
funders, media) over 40 monthly rounds. This release contains 119 LLM-mode runs (agent policies: claude-opus-4-6, claude-sonnet-4-6, gpt-5.5-2026-04-23) and 250 heuristic-mode runs… See the full description on the dataset page:
https://huggingface.co/datasets/evaluation-ecosystem/evaluation-ecosystem-data.