MixEval is a A dynamic benchmark evaluating LLMs using real-world user queries and benchmarks, achieving a 0.96 model ranking correlation with Chatbot Arena and costs around $0.6 to run using GPT-3.5 as a Judge.
You can find more information and access the MixEval leaderboard here.
This is a fork of the original MixEval repository. The original repository can be found here. I created this fork to make the integration and use of MixEval easier during the training… See the full description on the dataset page:
https://huggingface.co/datasets/zeitgeist-ai/mixeval.