Dataset Card for Evaluation run of Undi95/Mistral-11B-TestBench3
Dataset Summary
Dataset automatically created during the evaluation run of model Undi95/Mistral-11B-TestBench3 on the Open LLM Leaderboard.
The dataset is composed of 61 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Undi95__Mistral-11B-TestBench3.