BenchmarkCards is a standardized documentation dataset for large language model (LLM) benchmarks.
Each card summarizes key information about an LLM benchmark, including its objectives, methodology, data sources, targeted risks, limitations, and ethical considerations.
We gratefully thank all benchmark authors who provided feedback and approval for the BenchmarkCards in this repository. Your collaboration is essential… See the full description on the dataset page:
https://huggingface.co/datasets/ASokol/BenchmarkCards.