Benchmark-Focused MLT Dataset
Overview
Total Samples: 682,707
Train Samples: 680,662
Test Samples: 2,000 (MLT당 200개씩)
MLT Labels: 10
Math: GSM8K, MATH
Reasoning: BBH, ARC-Challenge
Knowledge: MMLU, MMLU-Pro
Commonsense: HellaSwag, Winogrande, PIQA
Truthfulness: TruthfulQA
Commonsense
96,413
14.1%… See the full description on the dataset page:
https://huggingface.co/datasets/Korea-MES/Mixtral-Upperbound.