This repository is dedicated to storing various evaluation data required for the glm-simple-evals evaluation project, to enable industry researchers and developers to reproduce the performance of the GLM-4.5 series models on reported benchmarks.
Currently, this repository covers the data required for the following evaluation tasks:
To use these evaluation datasets… See the full description on the dataset page:
https://huggingface.co/datasets/zai-org/glm-simple-evals-dataset.